Egy mérés, ami nemet mondottA measurement that said no
Eszközöket adtunk a munkásoknak, hogy pontosabban dolgozzanak. A mérés szerint ugyanannyi feladat sikerült, csak drágábban. Az eszközök alapból kikapcsolva maradtak.
We gave the workers tools to make them more accurate. The measurement showed the same number of successes, just at a higher price. The tools stay off by default.
Az ötletThe idea
Két eszközt építettünk, mindkettő ésszerűnek tűnt:
We built two tools, and both looked sensible:
- Önellenőrzés. A munkás vázlata még a kör vége előtt átmegy a szintaxis- és a build-kapun. Ha hibát talál, a munkás kap egy javító hívást, mielőtt a vázlat a közös kódba kerülne.
- Interfész-térkép. A munkás promptjába bekerül a projekt fájljainak külső felülete – útvonalak, exportok, függvények, űrlapmezők –, nagyjából 9000 karakterig. Így nem kell kitalálnia, mit kínál a szomszéd fájl.
- Self-check. A worker’s draft goes through the syntax and build gates before the round ends. If they find a problem, the worker gets one repair call before the draft joins the shared code.
- Interface map. The worker’s prompt includes the external surface of the project’s files – routes, exports, functions, form fields – up to roughly 9,000 characters. That way it doesn’t have to guess what the neighbouring file offers.
Az A/B mérésThe A/B measurement
Egyik ötlet sem volt légből kapott. Ha a munkás maga veszi észre a szintaxishibát, megspórolhat egy teljes kört; ha látja a szomszéd fájlok felületét, kevesebbet kell találgatnia. A kérdés csak az volt, hogy a gyakorlatban is így van-e – és ezt nem érzésre akartuk eldönteni.
Neither idea came out of nowhere. If a worker catches its own syntax error, it can save a whole round; if it sees the neighbouring files’ surface, it has less to guess. The only question was whether this holds in practice – and we didn’t want to decide that on gut feeling.
Ugyanazon a 16 futáson mértünk (8 feladat × 2): Flask, Node, React és bővítési feladatok vegyesen. Egyszer eszközök nélkül, egyszer eszközökkel.
We measured on the same 16 runs (8 tasks × 2): a mix of Flask, Node, React and extension tasks. Once without the tools, once with them.
A két változat között csak az eszközök kapcsolója különbözött: ugyanazok a feladatok, ugyanaz a független ellenőrző teszt, ugyanazok a modellek (a munkások Gemini 3.8 Flash, a kapuőr Gemini 3.6 Flash modellen futott).
The only difference between the two variants was the tools switch: same tasks, same independent check, same models (workers on Gemini 3.8 Flash, gatekeeper on Gemini 3.6 Flash).
| Változat | OK | Költség |
|---|---|---|
| Eszközök nélkül | 13/16 | 1,14 USD |
| Eszközökkel | 13/16 | 1,52 USD (+33%) |
| Variant | OK | Cost |
|---|---|---|
| Without tools | 13/16 | USD 1.14 |
| With tools | 13/16 | USD 1.52 (+33%) |
Harmadával több költség, pontosan ugyanannyi siker.
A third more cost, exactly the same number of successes.
Miért nem segített?Why didn’t it help?
Az önellenőrzés egyetlen egyszer sem sült el. A futásokat elrontó hibák teszthibák voltak, azokat pedig az önellenőrzés szándékosan nem nézi: egy bukott tesztet nem lehet megbízhatóan egyetlen munkáshoz rendelni, így nem is tudnánk, kinek adjuk a javító hívást. A bővítések ráadásul patch-módban futnak, így ott az önellenőrzés eleve nem lépett működésbe.
The self-check never fired, not once. The errors that sank runs were test failures, and the self-check deliberately ignores those: a failing test cannot be reliably attributed to one worker, so we wouldn’t know who should get the repair call. On top of that, extensions run in patch mode, so the self-check never came into play there.
Az interfész-térkép pedig csak a bemenetet növelte. Több token ment be minden hívásba, de a sikerek száma nem változott. Ez egyszerű számtan: amit minden hívásba beleteszünk, azt minden hívásnál ki is fizetjük.
The interface map, for its part, only made the input bigger. More tokens went into every call, and the success count didn’t move. Simple arithmetic: whatever you put into every call, you pay for on every call.
Amit a mérés helyette megmutatottWhat the measurement showed instead
A bukott futások okait végignézve kiderült, hogy a valódi hibák egészen máshol voltak:
Going through the failed runs, it turned out the real causes were somewhere else entirely:
- a modell érvénytelen JSON-t adott vissza – nyers sortörés, escape nélküli idézőjel a kódban (erről a Türelmes JSON cikkben írunk);
- egy React-hiba valódi vite build-hiba volt, amelyet a független teszt nem fedett le;
- a bővítéseknél a szerződés-törés: a csapat megváltoztatta a meglévő interfészt (lásd A meglévő teszt a szerződés).
- the model returned invalid JSON – raw newlines, unescaped quotes in code (see Tolerant JSON);
- one React failure was a genuine vite build error that the independent test didn’t cover;
- in extensions, contract breaks: the team changed the existing interface (see Existing tests are the contract).
Egyik ellen sem az önellenőrzés vagy az interfész-térkép a gyógyszer. A mérés tehát nemcsak nemet mondott, hanem megmutatta, hová érdemes a következő munkát tenni.
Neither the self-check nor the interface map is the cure for any of these. So the measurement didn’t just say no – it showed us where the next piece of work should go.
A döntésThe decision
Az eszközök alapból kikapcsolva maradnak. Munkásonként bekapcsolhatók a csapat-grafikon kártyáin, a parancssori változatban és a mérőrendszerben pedig kapcsolóval. Aki kísérletezni szeretne, megteheti – de alapértelmezésnek csak az kerül be, amit a mérés igazol.
The tools stay off by default. They can be enabled per worker on the team graph cards, and with a flag in the CLI and in the evaluation harness. Anyone who wants to experiment can – but only what the measurements back up becomes a default.
Tanulság. Egy funkció akkor jó, ha a mérés igazolja. Ha nem, marad kikapcsolva – akkor is, ha a fejlesztése sokáig tartott.
Takeaway. A feature is good when the measurements prove it. If they don’t, it stays off – however long it took to build.