Lab 6

De la WikiLabs
Jump to navigationJump to search

Lab 6 — Project Development II

Objectives

This laboratory continues the supervised development of the semester project.

The objective is to move from a working minimum viable product to a stable release candidate.

At this stage, the central functionality should already exist. The focus is therefore on completing missing features, improving reliability, removing technical debt, expanding tests, documenting the project and ensuring that another person can install and run it reproducibly.

After completing this laboratory, you should be able to:

  • identify and prioritize remaining work;
  • distinguish critical defects from optional improvements;
  • complete the central functionality of the project;
  • add regression tests for discovered bugs;
  • test important edge cases;
  • refactor code safely using tests;
  • improve logging and diagnostics;
  • improve project documentation;
  • validate installation from a clean environment;
  • verify dependency declarations;
  • document known limitations;
  • prepare a stable release candidate.

Most of the laboratory should be spent working directly on the project.

1. Laboratory structure

Time Activity
0–15 min Project status and remaining-work review
15–30 min Short discussion: release quality and regression testing
30–90 min Supervised implementation and refactoring
90–105 min Test and documentation review
105–120 min Release-candidate verification

2. From MVP to release candidate

The MVP demonstrated that the main technical idea works.

A release candidate should demonstrate that the project is close to final delivery.

prototype
    ↓
MVP
    ↓
feature complete
    ↓
tested
    ↓
documented
    ↓
release candidate
    ↓
final release

A release candidate does not need to be perfect, but no major known defect should prevent the primary workflow from operating correctly.

3. Feature completeness

Before adding new features, classify remaining work.

Category Meaning
required necessary for the project to satisfy its main objective
important significantly improves usability or robustness
optional useful enhancement but not required
experimental interesting idea that may not be stable enough for final integration

Required functionality should be completed before optional extensions.

4. Priority levels

A simple priority system is sufficient.

Priority Meaning
P0 application cannot perform its primary function
P1 major defect affecting normal use
P2 important improvement or non-critical defect
P3 optional improvement

P0 and P1 issues should be resolved before adding new optional features.

5. Remaining-work list

Maintain a short actionable list.

[ ] Handle malformed input
[ ] Add timeout handling
[ ] Add pipeline tests
[ ] Document configuration
[ ] Remove debug output
[ ] Verify clean installation

Avoid vague items such as:

[ ] Improve code
[ ] Fix things
[ ] Finish project

6. Stable software before more features

At this stage, a smaller stable project is generally preferable to a larger unstable one.

Prefer:

working feature A
working feature B
working feature C

over:

feature A mostly works
feature B sometimes works
feature C works
feature D unfinished
feature E experimental

7. Regression testing

When a defect is discovered:

discover bug
    ↓
reproduce bug
    ↓
write failing test
    ↓
fix implementation
    ↓
test passes
    ↓
keep test permanently

The resulting test is a regression test.

8. Regression test example

Suppose:

def average(values):
    return sum(values) / len(values)

The function fails incorrectly for an empty list.

Add a test:

import pytest


def test_average_rejects_empty_list():
    with pytest.raises(ValueError):
        average([])

Then update the function:

def average(values):
    if not values:
        raise ValueError(
            "values cannot be empty"
        )

    return sum(values) / len(values)

The test should remain in the project after the defect is fixed.

9. Edge-case review

Normal input is not sufficient.

For numerical data, consider:

minimum valid value
maximum valid value
zero
negative values
very large values
empty input
one-element input

For text:

empty string
whitespace-only string
Unicode characters
very long strings
unexpected encoding

For files:

missing file
empty file
malformed file
permission error
missing output directory
unexpected format

10. Network projects

Network projects should consider:

  • connection refused;
  • timeout;
  • server unavailable;
  • malformed response;
  • unexpected status code;
  • lost connection;
  • incomplete message.

Do not assume every network operation succeeds.

11. Hardware projects

Hardware projects should consider:

  • device unavailable;
  • device disconnected;
  • incomplete data;
  • invalid data;
  • timeout;
  • permission error;
  • unsupported device or firmware version.

Expected hardware failures should produce understandable application behavior.

12. Image-processing projects

Image-processing projects should consider:

  • unsupported format;
  • invalid dimensions;
  • empty image;
  • grayscale versus color input;
  • unexpected channel count;
  • corrupted input;
  • very large images.

Validate assumptions close to the input boundary.

13. Public behavior versus implementation detail

Tests should focus on public behavior.

Prefer:

result = process(image)

assert result.shape == expected_shape

over:

assert processor._temporary_buffer == ...

Private implementation details should remain free to change during refactoring.

14. Do not delete inconvenient tests

If a valid test fails after a code change, investigate the reason.

A test should normally be removed only when:

  • the requirement changed;
  • the tested behavior no longer exists;
  • the test itself is incorrect.

15. Test organization

A larger project can organize tests by component:

tests/
├── test_models.py
├── test_services.py
├── test_storage.py
├── test_protocol.py
└── test_pipeline.py

For larger projects:

tests/
├── unit/
│   ├── test_filters.py
│   └── test_protocol.py
└── integration/
    ├── test_storage.py
    └── test_pipeline.py

Use the simplest structure that remains understandable.

16. Unit versus integration tests

Example Type
validate one grade value unit
decode one packet unit
encode and then decode a packet integration
save and reload data from a temporary file integration
run a complete processing pipeline integration

Both kinds of tests are useful.

17. End-to-end testing

An end-to-end test exercises a complete workflow.

input
  ↓
application
  ↓
processing
  ↓
output

End-to-end tests are valuable but should not replace focused unit tests.

18. Testing CLI applications

Separate argument parsing from application logic.

Tightly coupled Easier to test
def main():
    args = parse_args()

    # All processing here.
    ...
def process_file(
    input_path,
    output_path,
):
    ...


def main():
    args = parse_args()

    process_file(
        args.input,
        args.output,
    )

The processing function can now be tested without command-line parsing.

19. Deterministic tests

Good tests produce the same result every time under the same conditions.

Avoid uncontrolled dependencies on:

  • current time;
  • random values;
  • network availability;
  • arbitrary filesystem state;
  • execution order.

For controlled randomness:

import random


rng = random.Random(1234)

20. Refactoring with tests

Use short cycles:

run tests
   ↓
all pass
   ↓
small refactor
   ↓
run tests
   ↓
all pass
   ↓
commit

Do not perform a very large refactoring before checking whether behavior is still correct.

21. Refactoring nested control flow

Before:

def process(record):
    if record is not None:
        if record.valid:
            if record.enabled:
                return transform(record)

    return None

Using guard clauses:

def process(record):
    if record is None:
        return None

    if not record.valid:
        return None

    if not record.enabled:
        return None

    return transform(record)

The second version reduces nesting.

22. Refactoring long conditions

Before:

if (
    device.connected
    and device.enabled
    and not device.error
    and device.temperature < 80
    and device.mode == "ready"
):
    ...

After:

def is_ready(device):
    return (
        device.connected
        and device.enabled
        and not device.error
        and device.temperature < 80
        and device.mode == "ready"
    )


if is_ready(device):
    ...

The function name expresses the meaning of the condition.

23. Structured data instead of loose dictionaries

For stable domain data, a dataclass may be clearer than an unstructured dictionary.

Dictionary Dataclass
student = {
    "name": "Alice",
    "grade": 9.5,
    "year": 2,
}
from dataclasses import dataclass


@dataclass
class Student:
    name: str
    grade: float
    year: int

Do not convert every dictionary into a class. Use the structure that best represents the data.

24. Avoid over-refactoring

Refactor when it improves:

  • clarity;
  • correctness;
  • maintainability;
  • testability;
  • reuse.

Do not introduce abstractions simply to make the architecture appear more complicated.

25. Logging review

Search for temporary debug output such as:

print("debug")
print(variable)
print("here")
print("test")

For each occurrence, decide whether it should be:

  • removed;
  • replaced with logging;
  • retained as intentional user output.

26. Logging levels

Level Typical use
DEBUG detailed internal processing information
INFO normal application events
WARNING unusual condition from which execution can continue
ERROR requested operation failed
CRITICAL application cannot continue

27. Useful logging context

Poor:

logger.error("Failed")

Better:

logger.error(
    "Failed to load configuration from %s",
    path,
)

Diagnostic messages should contain enough context to be useful.

28. Exception logging

Inside an exception handler:

try:
    load_device()
except DeviceError:
    logger.exception(
        "Device initialization failed"
    )
    raise

logger.exception() includes traceback information.

29. User-facing errors

Expected errors should usually produce concise user-facing messages.

For example:

Error: input file 'data.json' does not exist.

A predictable user error does not normally require displaying an internal traceback.

30. Documentation layers

Documentation Audience
README user or developer starting the project
docstrings developer using functions and classes
comments developer understanding non-obvious implementation decisions
release notes user/evaluator understanding changes and limitations

31. README checklist

The README should include:

  • project name;
  • short description;
  • main features;
  • requirements;
  • installation;
  • configuration;
  • how to run;
  • example usage;
  • how to run tests;
  • high-level architecture;
  • known limitations.

32. Reproducible installation

Installation instructions should be explicit.

git clone ...
cd project

python -m venv .venv
source .venv/bin/activate

python -m pip install -e .

Windows PowerShell:

.venv\Scripts\Activate.ps1

33. Running instructions

Document the actual command.

For example:

python -m project_name

or:

python -m project_name \
    --input data/input.json \
    --output results/output.json

34. Document configuration

If a project uses:

{
    "server": "localhost",
    "port": 8000,
    "timeout": 5.0
}

document every field.

Field Meaning
server remote server hostname
port remote service port
timeout request timeout in seconds

35. Environment variables

Do not document or commit real secrets.

Document variable names:

API_KEY
SERVER_URL
LOG_LEVEL

Example:

export API_KEY="..."

36. Docstrings

Use docstrings when behavior is not obvious.

def normalize(
    values: list[float],
) -> list[float]:
    '''Scale values to the interval [0, 1].

    Raises:
        ValueError: If all values are equal.
    '''
    ...

Do not duplicate obvious type information already expressed by type hints.

37. Public versus internal API

Internal names can begin with an underscore.

def _parse_header(data):
    ...


def load_frame(path):
    ...

This indicates that _parse_header() is an implementation detail.

38. Keep the public API small

Prefer a small meaningful interface:

load_image()
process_image()
save_image()

rather than exposing every internal intermediate operation.

39. Clean-environment verification

Do not rely on globally installed packages.

A useful verification is:

python -m venv test-env
source test-env/bin/activate

python -m pip install -e .
pytest

If this fails, installation instructions or dependency declarations are incomplete.

40. Verify declared dependencies

If code contains:

import requests
import numpy

the required packages must be declared in project metadata.

Another machine should not depend on packages that happened to be installed on the developer's system.

41. Remove unused dependencies

If the project declares:

dependencies = [
    "requests",
    "numpy",
    "pandas",
    "matplotlib",
    "opencv-python",
]

but only uses requests and numpy, remove unnecessary packages.

42. Version constraints

Avoid unnecessarily strict pinning unless required.

Possibly too strict:

dependencies = [
    "requests==2.32.4",
]

Potentially more flexible:

dependencies = [
    "requests>=2.32,<3",
]

The exact constraint depends on project compatibility requirements.

43. Supported Python version

Declare the expected Python version:

[project]
requires-python = ">=3.11"

Do not claim compatibility with an older version while using unsupported newer syntax.

44. Clean repository

The repository should generally exclude generated content such as:

.venv/
__pycache__/
*.pyc
.coverage
.pytest_cache/
.mypy_cache/
.ruff_cache/
build/
dist/

45. Example .gitignore

.venv/
__pycache__/
*.pyc

.pytest_cache/
.mypy_cache/
.ruff_cache/

.coverage
htmlcov/

build/
dist/
*.egg-info/

Add project-specific generated files where appropriate.

46. Sample data

If the project requires sample input for demonstration, include a small suitable example.

examples/
├── sample_input.json
└── expected_output.json

Do not rely on private local data that cannot be reproduced.

47. Sample configuration

For machine-specific configuration, provide:

config.example.json

rather than committing real credentials in:

config.json

48. Security review

Check for accidental secrets:

  • passwords;
  • API keys;
  • private keys;
  • access tokens;
  • database credentials;
  • credential-bearing URLs.

If a credential was committed, removing it from the latest version may not be sufficient. The credential should be rotated.

49. Input validation review

Review every external input boundary:

  • command-line arguments;
  • configuration;
  • files;
  • network messages;
  • user input;
  • hardware data;
  • API responses.

Ask:

What assumptions does this code make?

Validate important assumptions before processing the data.

50. Serialization round trip

If data is stored and later loaded, test a round trip.

object
  ↓
serialize
  ↓
file
  ↓
deserialize
  ↓
object

Example:

def test_round_trip(tmp_path):
    path = tmp_path / "data.json"

    original = [
        Student("Alice", 9.5),
        Student("Bob", 8.0),
    ]

    save_students(path, original)

    restored = load_students(path)

    assert restored == original

51. Performance sanity check

Full optimization is not required, but the application should be tested on representative input.

If the intended workload is:

100000 records
100 images
continuous packets
large matrices

do not test only with one tiny input and assume performance is acceptable.

52. Resource review

Check for:

  • files left open;
  • sockets not closed;
  • unnecessary large copies;
  • unbounded queues/lists;
  • repeated loading of large data;
  • infinite retry loops;
  • threads or processes not terminated.

Use context managers where appropriate.

53. Graceful shutdown

Applications managing external resources should stop predictably.

Example:

try:
    application.run()
finally:
    application.close()

When appropriate, make resources context managers.

54. Keyboard interrupt

A CLI application may handle Ctrl+C cleanly:

def main():
    try:
        run()
    except KeyboardInterrupt:
        print("\nInterrupted.")

Only add this behavior when it improves the application.

55. Release version

The project should have a version identifier.

[project]
version = "0.9.0"

A release candidate can be described as:

0.9.0-rc1

56. Release notes

A short release note can contain:

Version
Main features
Important fixes
Known limitations
Installation notes

It does not need to be long.

57. Known limitations

Document real limitations explicitly.

Examples:

Only PNG input is supported.
IPv6 is not supported.
The program has only been tested on Linux.
Maximum tested image size is 4096×4096.
Device reconnect requires restart.

58. Known issue versus critical defect

A known issue may remain if:

  • it does not prevent the primary workflow;
  • its scope is clear;
  • fixing it immediately would create disproportionate risk.

A critical defect should not simply be renamed a limitation.

59. Manual acceptance test

In addition to automated tests, perform a manual end-to-end test.

[ ] clean startup
[ ] representative input accepted
[ ] main feature works
[ ] output correct
[ ] expected error handled
[ ] application exits cleanly

60. Clean-machine test

Ask:

Could another student clone this repository
and run the project using only the README?

Test it:

  1. create a clean environment;
  2. install dependencies;
  3. run tests;
  4. start the application;
  5. follow only documented steps.

61. Project review checklist

Area Question
features Are all required features implemented?
defects Are any P0 or P1 defects open?
tests Are central behavior and major failures tested?
architecture Is remaining technical debt acceptable?
logging Are diagnostics useful and appropriately leveled?
errors Are expected failures handled?
documentation Can another user install and run the project?
dependencies Are all required packages declared?
repository Is generated and sensitive content excluded?
release Can the application be demonstrated reliably?

62. Exercise 1 — Remaining-work triage

List all remaining tasks.

Assign:

P0
P1
P2
P3

Resolve P0 and P1 issues before optional work.

63. Exercise 2 — Regression test

Choose one defect discovered since Lab 5.

Then:

  1. reproduce the defect;
  2. add a failing automated test;
  3. fix the implementation;
  4. confirm the new test passes;
  5. run the complete test suite.

64. Exercise 3 — Edge-case audit

Choose the most important input to the project.

Document at least five edge cases.

Input Expected behavior
empty input clear validation error or defined empty result
valid minimal input processed successfully
malformed input parsing error reported
missing input missing-resource error
large representative input completes within reasonable resources

Add missing tests where practical.

65. Exercise 4 — Test-suite cleanup

Review the test suite for:

  • duplicated setup;
  • unclear names;
  • order dependencies;
  • permanent local test files;
  • unnecessary real network access;
  • missing assertions;
  • disabled tests.

Refactor where appropriate.

66. Exercise 5 — Refactoring checkpoint

Choose one area with technical debt:

  • long function;
  • large class;
  • duplicated logic;
  • deep nesting;
  • hard-coded configuration;
  • global state.

Use:

tests pass
   ↓
small refactor
   ↓
tests pass
   ↓
commit

67. Exercise 6 — Logging audit

Search for:

print(
logger.debug(
logger.info(
logger.warning(
logger.error(
logger.exception(

Check that:

  • user output intentionally uses print();
  • diagnostics use logging;
  • levels are appropriate;
  • errors include context;
  • secrets are never logged.

68. Exercise 7 — README installation test

Follow the README in a clean environment.

Record every missing step.

Update the README until no undocumented knowledge is required to start the project.

69. Exercise 8 — Dependency verification

Create a fresh virtual environment.

Install the project using its documented procedure.

Then run:

pytest
ruff check .

If used by the project:

mypy src

Fix dependency or configuration problems.

70. Exercise 9 — Repository audit

Run:

git status

Inspect the repository for:

  • virtual environments;
  • caches;
  • generated output;
  • secrets;
  • large temporary files;
  • local configuration;
  • editor-specific files.

Update .gitignore if required.

71. Exercise 10 — Known limitations

Add:

## Known limitations

to the README.

List only real limitations of the current implementation.

72. Exercise 11 — Release notes

Write:

Release candidate: 0.9

Implemented:
- ...

Fixed:
- ...

Known limitations:
- ...

Run:
- ...

73. Exercise 12 — Acceptance test

Execute the complete main workflow.

Record:

Input:
Command:
Expected output:
Actual output:
Result: PASS / FAIL

A failure in the primary workflow should be treated as a high-priority issue.

74. Release-candidate requirements

By the end of Lab 6:

Requirement Expected state
core functionality complete
P0 defects none known
P1 defects ideally none
tests central functionality and major failures covered
regression tests added for important discovered defects
documentation installation and usage complete
dependencies reproducible in a clean environment
logging temporary debug output removed
repository clean and free of secrets
limitations explicitly documented
demonstration repeatable end-to-end

75. Instructor review

Area Review question
stability Does the primary workflow work consistently?
completeness Are required features implemented?
testing Are important edge cases tested?
regression Were discovered defects converted into tests?
design Is the architecture still understandable?
documentation Can the project be installed without assistance?
robustness Are predictable failures handled?
release readiness Can the team demonstrate the project reliably?

76. C/C++ habits to reconsider

C/C++ habit Typical Python project approach
rely mainly on compile success verify tests, linting and runtime behavior
test manually only at the end maintain regression tests continuously
treat documentation as optional document installation, configuration and execution
assume the developer machine represents deployment verify in a clean virtual environment
leave debug output in code use structured logging
keep redesigning until submission stabilize architecture before final delivery
add features until the last moment freeze core features and focus on reliability

77. Feature freeze

After Lab 6, the project should enter a practical feature freeze.

This means:

  • no major architecture changes;
  • no unnecessary new dependencies;
  • no large experimental features;
  • focus on defects;
  • focus on documentation;
  • focus on final demonstration reliability.

Small low-risk improvements may still be made.

78. Final preparation order

A useful order is:

critical defects
     ↓
required features
     ↓
tests
     ↓
documentation
     ↓
clean installation
     ↓
demo preparation
     ↓
optional polish

Do not prioritize presentation polish while critical project defects remain.

79. Summary

Lab 6 is primarily a stabilization laboratory.

The project should move from:

"It basically works."

to:

"It can be installed,
tested,
run,
demonstrated,
and understood
by another person."

The central idea is:

A release candidate is not simply code with all planned features. It is software whose behavior, dependencies, limitations and execution procedure are sufficiently controlled to be demonstrated and evaluated reliably.

80. Preparation for Lab 7

Before the final laboratory:

  1. freeze the main feature set;
  2. fix remaining critical defects;
  3. run the complete automated test suite;
  4. perform a clean-environment installation;
  5. verify the README;
  6. document known limitations;
  7. prepare a stable demonstration dataset or scenario;
  8. ensure the demonstration does not depend on unavailable external resources;
  9. prepare a concise architecture explanation;
  10. prepare to explain one important design decision;
  11. prepare to explain one significant technical difficulty.

In Lab 7, each team will perform the final demonstration, code walkthrough and technical discussion.