A meglévő teszt a szerződésExisting tests are the contract
Egy meglévő projekt bővítésénél a csapat átírta a régi teszteket, hogy zöldek legyenek. Megjavítottuk – és a következő mérés rosszabbnak tűnt. Valójában őszintébb volt.
While extending an existing project, the team rewrote the old tests until they went green. We fixed that – and the next measurement looked worse. In fact it was more honest.
A feladatThe task
Két bővítési feladatot mértünk: egy meglévő Flask jegyzet-API-t kellett címkékkel kiegészíteni (brown-notes-tags), és egy többmodulos alkalmazást ugyanígy (brown-multi-tags). A cél mindkét esetben kimondta: a meglévő viselkedés nem változhat, a meglévő teszteket nem szabad törölni vagy gyengíteni.
We measured two extension tasks: adding tags to an existing Flask notes API (brown-notes-tags), and doing the same to a multi-module app (brown-multi-tags). In both cases the goal said it plainly: existing behaviour must not change, and existing tests must not be deleted or weakened.
1. mérés: egy hamis zöld és a történeteRound 1: one false green, and how it happened
Négy futásból három volt OK, egy HAMIS ZÖLD. Ez utóbbit érdemes lépésről lépésre végigkövetni. Az Architect terve megváltoztatta a GET /api/notes válaszát:
Out of four runs, three were OK and one was a FALSE GREEN. That one is worth following step by step. The Architect’s plan changed the response of GET /api/notes:
# szemléltetés – eredetileg:
{"notes": [{"id": 1, "body": "..."}]}
# a terv szerint:
[{"id": 1, "content": "..."}]
# illustration – originally:
{"notes": [{"id": 1, "body": "..."}]}
# according to the plan:
[{"id": 1, "content": "..."}]
A meglévő tesztek ezt az első körben helyesen jelezték: három bukás. Erre a tesztíró munkás átírta a meglévő teszteket az új alakra, az egyik ellenőrzést pedig lazította – a body vagy a content is jó lett. A kapuk kizöldültek, a kapuőr átengedte. A független teszt 10-ből 8 bukást mutatott. A Backend közben írt egy osztályt is, amely a régi indexelést „hamisította” volna – ez halott kód maradt.
The existing tests caught this correctly in the first round: three failures. The test writer then rewrote those existing tests to match the new shape and loosened one check – either body or content would now do. The gates went green and the gatekeeper let it through. The independent test showed 8 failures out of 10. Meanwhile the Backend had written a class that would have “faked” the old indexing – it ended up as dead code.
1. javítás: a meglévő tesztek védelmeFix 1: protecting existing tests
A futás előtt zöld tesztfájlokat elmentjük. Ha a csapat átírja vagy törli őket, az eredeti változatuk is lefut az új kódon, és ha az bukik, a tesztkapu piros. A védelem kikapcsolható arra az esetre, ha valaki szándékosan akar viselkedést változtatni.
Test files that were green before the run are saved. If the team rewrites or deletes them, the original versions still run against the new code, and if they fail, the test gate is red. The protection can be switched off for when someone deliberately wants to change behaviour.
2. mérés: rosszabbnak tűntRound 2: it looked worse
Négy futás: 1 OK, 0 hamis zöld, 2 FELESLEGES PIROS, 1 BUKOTT. Elsőre visszalépésnek látszott. De a független teszt szerint a kód ugyanolyan jó volt, mint korábban (4-ből 3 helyes) – csak most az FRST nem hazudott róla. Az egyik BUKOTT futásnál a védelem éppen egy újabb hamis zöldet fogott meg: ott a független teszt 7 bukást mutatott.
Four runs: 1 OK, 0 false green, 2 UNNECESSARY RED, 1 FAILED. At first glance, a step backwards. But according to the independent test the code was just as good as before (3 out of 4 correct) – FRST simply wasn’t lying about it any more. In the FAILED run the protection had caught another would-be false green: the independent test showed 7 failures there.
A futásnaplókból kiderült a közös ok. Az Architect csak a célt kapta meg, a meglévő kódból egyetlen sort sem látott, ezért kitalálta az interfészt. Mindhárom rossz futás terve átírta a meglévő API-t. A kapuőr pedig a tervhez mérte a kódot, és elutasította a helyes, a meglévő felületet megőrző megoldást azzal, hogy „a PROJECT_PLAN.md szerint JSON listát kell visszaadjon”.
The run logs revealed the common cause. The Architect only received the goal; it never saw a single line of the existing code, so it invented the interface. The plans of all three bad runs rewrote the existing API. The gatekeeper then judged the code against the plan and rejected the correct solution – the one that preserved the existing interface – because “according to PROJECT_PLAN.md it must return a JSON list”.
A hiba nem a tesztíró munkásnál volt, hanem ott, ahol a terv készült: vakon.
The fault wasn’t with the test writer. It was where the plan was made: blind.
2. javítás: a terv lássa a szerződéstFix 2: the plan sees the contract
Meglévő projektnél az Architect ma látja a meglévő kódot, és szó szerint a zöld teszteket. A terv végére a rendszer egy „SZERZŐDÉS” szakaszt ír: ellentmondás esetén a meglévő teszt nyer, a terv téved. A kapuőr ugyanezt a szabályt kapja.
For an existing project the Architect now sees the existing code and, verbatim, the green tests. The system appends a “CONTRACT” section to the plan: in a conflict the existing test wins and the plan is wrong. The gatekeeper gets the same rule.
A három mérés egymás mellettThree rounds side by side
| Mérés | Futás | OK | HAMIS ZÖLD | FELESLEGES PIROS | BUKOTT |
|---|---|---|---|---|---|
| 1. – védelem nélkül | 4 | 3 | 1 | 0 | 0 |
| 2. – tesztvédelem | 4 | 1 | 0 | 2 | 1 |
| 3. – tesztvédelem + szerződés a tervben | 6 | 6 | 0 | 0 | 0 |
| Round | Runs | OK | FALSE GREEN | UNNECESSARY RED | FAILED |
|---|---|---|---|---|---|
| 1 – no protection | 4 | 3 | 1 | 0 | 0 |
| 2 – test protection | 4 | 1 | 0 | 2 | 1 |
| 3 – test protection + contract in the plan | 6 | 6 | 0 | 0 | 0 |
A 3. mérés 2 feladat × 3 futás volt: 6/6 OK, egyetlen hamis zöld nélkül, mindegyik futás 1 kör alatt kész, összesen 0,46 USD-ért – futásonként nagyjából 0,08 USD.
Round 3 was 2 tasks × 3 runs: 6/6 OK with no false green, every run finished in a single round, for USD 0.46 in total – roughly USD 0.08 per run.
Tanulság. A mérés, amely rosszabbnak tűnt, valójában őszintébb volt – és megmutatta a valódi hibát. A meglévő teszt nem lehet a csapat kezében.
Takeaway. The round that looked worse was simply more honest – and it pointed to the real bug. Existing tests must not be in the team’s hands.