measured
Research
Studies I ran on my own work, published with the numbers behind them. Newest first. Every one is a PDF you can take away, and one of them is a retraction of an earlier finding of mine.
-
Do the cheap agents pay for themselves? Seven delegations, measured
do the cheap agents pay for themselves: seven instrumented delegations from one session; 3 of 7 caught something I had missed, 1 was a false finding, and no saving is claimed because the counterfactual is not measurable
-
The translation audit: a local model re-reads my Finnish
the translation audit: a local Finnish model re-reads all 396 of the site's Finnish strings against their English source; only 2 of its 276 proposed rewrites held up
-
Poro-2-8B in production: what we measured, what broke, what we built around it
Poro-2-8B in production: what two projects measured, why one adopted it and one passed, and the deterministic layer built around it
-
Which local model writes the best Finnish? A blind test settles it
the blind test: a native speaker ranks 3 local models blind on Finnish naturalness; Poro wins 26/30
-
How a Finnish-RAG experiment caught and corrected its own mistake
finnish rag, the methodology: how the experiment caught and corrected its own mistake; the process, not the findings
-
Does the portfolio RAG need Finnish, and does it need a Finnish-built model?
the rag finnish experiment: 3 local 8B models on Finnish synthesis vs containment, single-variable, €0
-
Skill-suite calibration: 96 A/B arms across three codebases and three models
latest + broadest: 16 skills, cold-vs-skill A/B across 3 models; the current snapshot
-
The two skill-auditors: what they cost, and what they fixed
the synthesis: what the two skill-auditors cost (~36% cheaper to run) and the traps they exposed
-
Round 6 of the skills optimization study: re-measuring the six noisiest cells
round 6 · the noisiest cells re-measured at depth: an N=1 fluke overturned, ~+76% confirmed
-
Why "read each SKILL.md" costs tokens: five rounds of before/after testing
the optimization: 5 rounds of before/after on a SKILL.md; 3 cost-traps found + fixed
-
The skill catalog: every skill across four repositories
every skill across all 4 repos: the inventory, with measured (not guessed) costs