Skip to content

measured

Research

Studies I ran on my own work, published with the numbers behind them. Newest first. Every one is a PDF you can take away, and one of them is a retraction of an earlier finding of mine.

  1. Do the cheap agents pay for themselves? Seven delegations, measured

    do the cheap agents pay for themselves: seven instrumented delegations from one session; 3 of 7 caught something I had missed, 1 was a false finding, and no saving is claimed because the counterfactual is not measurable

    Read PDF · 142 KB

  2. The translation audit: a local model re-reads my Finnish

    the translation audit: a local Finnish model re-reads all 396 of the site's Finnish strings against their English source; only 2 of its 276 proposed rewrites held up

    About PDF · 180 KB

  3. Poro-2-8B in production: what we measured, what broke, what we built around it

    Poro-2-8B in production: what two projects measured, why one adopted it and one passed, and the deterministic layer built around it

    Read PDF · 129 KB

  4. Which local model writes the best Finnish? A blind test settles it

    the blind test: a native speaker ranks 3 local models blind on Finnish naturalness; Poro wins 26/30

    About PDF · 228 KB

  5. How a Finnish-RAG experiment caught and corrected its own mistake

    finnish rag, the methodology: how the experiment caught and corrected its own mistake; the process, not the findings

    Read PDF · 128 KB

  6. Does the portfolio RAG need Finnish, and does it need a Finnish-built model?

    the rag finnish experiment: 3 local 8B models on Finnish synthesis vs containment, single-variable, €0

    About PDF · 265 KB

  7. Skill-suite calibration: 96 A/B arms across three codebases and three models

    latest + broadest: 16 skills, cold-vs-skill A/B across 3 models; the current snapshot

    Read PDF · 424 KB

  8. The two skill-auditors: what they cost, and what they fixed

    the synthesis: what the two skill-auditors cost (~36% cheaper to run) and the traps they exposed

    Read PDF · 234 KB

  9. Round 6 of the skills optimization study: re-measuring the six noisiest cells

    round 6 · the noisiest cells re-measured at depth: an N=1 fluke overturned, ~+76% confirmed

    Read PDF · 140 KB

  10. Why "read each SKILL.md" costs tokens: five rounds of before/after testing

    the optimization: 5 rounds of before/after on a SKILL.md; 3 cost-traps found + fixed

    Read PDF · 352 KB

  11. The skill catalog: every skill across four repositories

    every skill across all 4 repos: the inventory, with measured (not guessed) costs

    About PDF · 375 KB