Célzott visszajelzésTargeted feedback
Egy bukott teszt hírét régen mindenki megkapta a csapatban. Ez pénzbe került, és új hibákat szült. Most a hiba annak megy, aki javítani tudja.
A failing test used to be broadcast to the whole team. That cost money and bred new bugs. Now the failure goes to whoever can actually fix it.
A probléma: mindenki javítThe problem: everyone fixes
Az FRST körökben dolgozik: a munkások írnak, a kapuk ellenőriznek, és ha valami piros, a hibaüzenet visszakerül a munkásokhoz a következő körre. Korábban egy bukott teszt visszajelzése mindenkihez ment, aki a körben bármilyen fájlt írt.
FRST works in rounds: workers write, gates check, and if something is red, the error goes back to the workers for the next round. Until recently, a failing test’s feedback went to everyone who had written any file in that round.
A mérésekben ez így nézett ki: egy backend-teszt bukását a Frontend is megkapta, pedig ő csak sablont és stílust írt. Megpróbált segíteni, és a „javítását” a main.py-ba írta – ütközés a Backenddel. Körről körre egy teljes, fizetős modellhívás ment el erre.
In our measurements it looked like this: a backend test failure also reached the Frontend worker, even though it had only written templates and styles. It tried to help and wrote its “fix” into main.py – colliding with the Backend. Every round, one full paid model call was spent on that.
Minden fölösleges hívás pénz, és minden fölösleges „javítás” újabb hibaforrás.
Every unnecessary call costs money, and every unnecessary “fix” is another source of bugs.
A megoldás: ki kapja a hibát?The fix: who gets the failure?
Első ránézésre a „mindenki kapja meg, hátha segít” biztonságos választásnak tűnik: ha több szem látja a hibát, nagyobb az esély, hogy valaki kijavítja. Egy AI-munkás azonban nem mond nemet egy hibaüzenetre. Ha megkapja, cselekedni fog – akkor is, ha a hiba nem az ő területén van. Így a széles körű értesítés nem több segítséget hozott, hanem több, egymást keresztező módosítást.
At first glance, “send it to everyone, just in case” looks like the safe choice: more eyes on the error, more chances someone fixes it. But an AI worker doesn’t shrug off an error message. If it receives one, it acts – even when the problem is outside its own files. So broadcasting didn’t bring more help; it brought more edits crossing each other.
Ma egy teszt bukása három helyre megy:
Today a test failure is routed to three kinds of recipient:
- a logikát író munkásokhoz – vagyis akiknek nem teszt-, nem dokumentáció- és nem megjelenítési (html, css, sablon, statikus) fájljuk van;
- ahhoz, akit a hibaszöveg megnevez (például az ő fájlja szerepel a hibában);
- a teszt gazdájához, azzal a kiegészítéssel, hogy a tesztet csak akkor módosítsa, ha bizonyíthatóan rosszat vár el.
- workers who write logic – that is, whose files are not tests, not documentation and not presentation (html, css, templates, static assets);
- whoever the error text names (for example, because their file appears in it);
- the owner of the test, with a note to change the test only if it provably expects the wrong thing.
Az utolsó kiegészítés fontos: a tesztíró ne az legyen, aki a pirosat zöldre festi. Erről szól A meglévő teszt a szerződés című cikkünk is.
That last note matters: the test writer should not be the one who paints red into green. Our article Existing tests are the contract is about exactly that.
Második probléma: a munkás, akinek nincs dolgaSecond problem: the worker with nothing to do
Egy bővítésnél a Documentation Writer minden körben újra meghívódott, pedig nem volt dolga – és minden körben ezt is válaszolta. A szabály most egyszerű: aki kétszer egymás után azt feleli, hogy nincs dolga, azt nem hívjuk újra, hacsak nem kap neki címzett visszajelzést.
In one extension task the Documentation Writer was called again every round although it had nothing to do – and said so every round. The rule is now simple: a worker that answers “nothing to do” twice in a row is not called again unless feedback is addressed to it.
Harmadik: aki nem vesz résztThird: workers who sit it out
Az Architect a tervben megjelölheti, ki nem vesz részt a munkában („NEM VESZ RÉSZT” sor). Az ilyen munkást meg sem hívjuk. Egy mérés azonban megmutatta, hogy ez nem volt elég: egy általános figyelmeztetés („még egyetlen fájlt sem készítettél”) felülírta a tervet, és a kihagyott Frontend és Documentation Writer a második körben mégis dolgozott. Ezt is javítottuk, és a regressziós visszajelzés is csak a logikát író munkásokhoz megy.
The Architect can mark workers in the plan as not taking part (a “NOT PARTICIPATING” line). Such a worker is not called at all. One measurement showed this wasn’t enough: a generic warning (“you haven’t produced a single file yet”) overrode the plan, and the excluded Frontend and Documentation Writer went to work in round two anyway. We fixed that too, and regression feedback now also goes only to the logic writers.
Eredmény – óvatosanResults – with caution
| Bővítési futások | Költség futásonként |
|---|---|
| Korábban | 0,13–0,30 USD |
| Célzott visszajelzéssel | kb. 0,11–0,13 USD |
| Legutóbbi mérés (1 kör) | kb. 0,08 USD |
| Extension runs | Cost per run |
|---|---|
| Before | USD 0.13–0.30 |
| With targeted feedback | approx. USD 0.11–0.13 |
| Latest measurement (1 round) | approx. USD 0.08 |
Ezek kevés futásból származó számok, ezért óvatosan kezeljük őket. Az irány viszont egyértelmű: kevesebb hívás, kevesebb ütközés, kevesebb kör.
These figures come from a small number of runs, so we treat them with caution. The direction is clear, though: fewer calls, fewer collisions, fewer rounds.
Tanulság. Minden fölösleges hívás pénz, és minden fölösleges „javítás” újabb hibaforrás. A hiba annak menjen, aki javítani tudja.
Takeaway. Every unnecessary call costs money, and every unnecessary “fix” adds risk. Send the failure to whoever can fix it.