Joshua DiniegaFull-Stack Software Engineer

What it is

I built Scar around one habit: every bug diagnosed in a session leaves a typed record, written by the agent that hit it. A pipeline clusters those records by shared mechanism, distils each cluster into one lesson, and serves the survivors to a coding agent over MCP. That is the half an agent queries. The other half runs in the agent's harness whether anything queries or not: gates that fire on a step not taken, and a scanner that greps every edit for the exact forms past bugs took. Agents report back what each lesson did, and that moves its ranking. The site is the distilled layer; its recall page runs the real keyword index in your browser.

Why I built it

I started the habit: write down every bug you diagnose. After a thousand-odd records the archive had gone write-only, and I kept hitting problems it already had the answer to.

Writing down every bug I diagnosed is part of how I work. After roughly a thousand records, most of them out of professional work, the archive had gone write-only. Comprehensive, never read, while the same class of mistake kept recurring. Handing the pile to a coding agent doesn't fix it either: agents already have repository search, git history and long context, and none of that lets them open a million words to find out they have hit this bug before. The gap is the knowledge nobody writes into a repository. We tried that before. That looks correct but breaks on native. That fix caused a regression last month. So the records stopped being the product and became a compiler's input. And the compiled form that matters most is a check: retrieval only helps the session that thinks to ask, so the lessons that name a grep-able form promote into a scanner and a set of harness gates that fire without being asked. The agents write the records, and report back what each lesson did — which leaves the question I actually care about: what stops a corpus that writes itself from filling up with confident nonsense? I don't have a complete answer. The gates below are what I have so far.

In the app

Captured from the public site, which is the distilled layer only — every source record behind it stays on my machine. The client-work projects appear in these frames as counts with their names masked, which is the whole demonstration.

By the numbers

source records compiled
1,299
distilled lessons
322
tokens an agent carries per session
~2.2K
lines of pipeline, tooling and viewer
89,873

Counted against the source on 29 September 2026, not estimated.

Decisions that mattered

14 of them. Each opens onto what I did and the reasoning behind it.

  1. The source records are never loaded — only what compiles out of them

    What I did, and why

    I never load the source records. A deterministic pass clusters them by shared error signature, symbol or vocabulary, and each cluster is rewritten as one lesson stated from the mechanism, not the incident. The cluster's size becomes its recurrence count. The raw records stay on disk as an audit trail, reachable only from the lesson they produced.

    When the reader has a context window, storing more and searching better is the wrong axis. What I care about is the ratio between what the corpus knows and what it costs to consult, and only a build step moves it. Rewriting from the mechanism is what lets a lesson transfer to a codebase that shares nothing with the one that produced it. The recurrence count does something narrower: it tells me whether a failure mode is real or I just had one bad afternoon.

  2. A lesson that names a grep-able form stops being prose

    What I did, and why

    I set knowledge up as a ladder: observation, lesson, mechanism. A lesson promotes once it names a form you can grep for, and it promotes out of the index entirely, into a scanner that runs on every edit unattended and harness gates that fire on a step not taken. Recurrence alone keeps it as prose, which is what recall ranks.

    Retrieval is pull-only. It helps the session that thinks to ask, and the sessions that most need a lesson are the ones with no reason to query for it. A check runs either way. The split also settled what prose is for: judgment-shaped lessons stay in the index for an agent to weigh, and the machine-checkable ones move somewhere stricter and cheaper than a sentence. The reframe reached the whole system in one dated pass. The front page stopped saying memory the same day, because a misframed bootstrap teaches every session the wrong model of what it's holding.

  3. Keyword ranking decides; a trained encoder may only reorder what it finds

    What I did, and why

    I kept BM25 as the base: keyword ranking over the lesson layer, with each lesson's declared triggers weighted three times its body, and the caller's own phrasings fused. Since 2026-09-28 a small encoder ranks the same lessons too, a MiniLM model fine-tuned on task-to-lesson pairs, 23 MB as int8. Its ranking is fused in only when keyword search already returned something. It never answers alone, and when the model is missing or fails its checksum, recall falls back to keywords and logs the fault.

    Until that date this decision said no embedding model at all, and my reasons still hold: every lesson states when it should resurface, and a caller that's itself a language model can phrase a task several ways for free. What changed was a measurement. The bar was written down before any training: beat keyword recall by ten points on lessons the model never saw, at p < 0.05. The first encoder missed it by 1.2 points, so it stayed off. The second, trained on 1,403 tasks, found the right lesson in the top three 65.0% of the time against keyword recall's 52.5%, on 80 held-out lessons, p = 0.003. I gated it because its similarity can't tell an off-topic query from an on-topic one. Ungated it scored higher, and it also answered questions the corpus has nothing for.

  4. When a query names files, recall reads the code and derives a phrasing of its own

    What I did, and why

    A recall that carries file paths gets one extra query variant for free, derived deterministically from the code itself — imports, declared names, called members, comments stripped — and fused with the caller's phrasings. Recorded outcomes decide which derived terms keep contributing: a term that only ever arrived with noise lands on a stoplist, and results that arrived only through the bridge are marked so the attribution survives into the ledger.

    The caller describes a task in their vocabulary and the corpus wrote its lessons in another, and that gap is exactly where keyword ranking loses to embeddings. Reading the file closes it without a model: the code already states its own nouns, so nobody has to guess the corpus's wording. Making the bridge earn itself from outcomes is the same rule I hold the rest of the ranking to: nothing the writer asserts is an input, only what happened in use.

  5. Client work is publishable as scale, never as content

    What I did, and why

    Records from employer or client work carry visibility: private, under a rule I set. The public build reduces each to a stub before it can reach a page: opaque id, a title rebuilt from schema fields, no body, no tags, no links. Only the character count survives, and two denylist tiers enforce that at build time.

    I drop everything identifying instead of masking it. A mask claims you enumerated all the identifiers, and free-text tags had already leaked a person's name once. Keeping the counts is deliberate: the site can say how much professional work sits behind the lessons without any of it being readable. The denylist catches names, not context, so it gates the build and doesn't replace a human read.

  6. A build-time preflight, because a bundler almost carried the corpus into a deploy

    What I did, and why

    I moved the check that the corpus exists out of the content layer into a preflight script that runs before the build. It also fails when the corpus is present but empty.

    A filesystem call on a computed path at module scope defeats static analysis, so Turbopack conservatively traced the whole repository into the server output, and this repository contains the private client records. Nothing was actually bundled; the trace manifests came back with zero markdown. But the failure mode is silent and the blast radius is client confidentiality, so I didn't let it stay by luck. The empty-corpus case is subtler: the build would otherwise succeed and publish a site whose every number is zero, looking intentional.

  7. The system's own running cost is a failing test, not a promise

    What I did, and why

    I put token ceilings on the two files that load into every session — the bootstrap every adopting project imports, and the kernel served once per session — and enforced them in the verify suite. Exceeding one fails verify. Raising a ceiling is allowed, but only as a diff in that script.

    The whole claim is that over a million words of experience cost a few thousand tokens a session, and only half of it was enforced: the kernel had a cap, the bootstrap file didn't, and it drifted from roughly 800 to 1,900 tokens across three rewrites without anyone deciding that. A knowledge system that's expensive to consult is worse than none, because it taxes every task whether or not it helps. A number stated in prose decays; the same number in a check that fails holds.

  8. The ledger prices what Scar spends, and stops at break-even

    What I did, and why

    Every token Scar puts into a context is priced in a committed ledger, split into eras so an experiment's spend accrues separately from the engine it changed and a frozen release keeps its own caveat. The measured side stops at break-even: anything past it is labelled a projection.

    The cost a lesson avoided is unobservable; only the cost it added can be measured, so an honest ledger is one-sided and says so on the page. The era split exists because a reverted experiment would otherwise smear its spend into the engine's baseline. A revert on record is the reason: an attempt to halve the always-loaded kernel was undone the same day, once it was clear the experiment couldn't answer its own question. The attempt, the revert and the reasoning are all in the decision log, which is the system doing what it claims: the next session argues with the record instead of re-running the experiment.

  9. Retrieval quality is a benchmark, because ranking cannot be eyeballed

    What I did, and why

    The verify suite runs task-phrased queries with known-correct answers and fails below an 80% top-three hit rate. The queries deliberately avoid the vocabulary the lessons declare as their own triggers. A second set of clearly off-domain queries (Kubernetes, Postgres tuning, payroll) is scored the same way and must return nothing.

    A tokenizer tweak that fixes the query in front of you silently breaks two you aren't looking at, and the failure surfaces later as an agent giving a confident wrong answer with no sign that retrieval caused it. It earned itself immediately: the baseline was 9 of 12, and the fix was a change to word-stem folding that would've been indistinguishable from a regression without a score. The negative set was added after measuring the false-positive rate and finding it wasn't zero: a coverage score counting query words present anywhere in the corpus let an off-domain query reach 'strong confidence' on one coincidental word. Matching a lesson by the words it chose for itself proves nothing, which is why the positive queries avoid them.

  10. The agents write the knowledge base, so nothing takes an agent's word for its own work

    What I did, and why

    Agents write the records, distil them into lessons and report what each lesson did. So I made every input to the ranking something the writing agent can't assert: how many independent records shared a mechanism, whether the claim names a form you can grep for, and what happened when it was used. The one self-graded field was a lesson rating its own confidence high, which multiplied its rank by up to two and a half. I made it earned instead of claimed, capped at write time and audited again over every file already on disk.

    A knowledge base that writes itself is only worth having if it can't flatter itself, and a self-assessed confidence score is exactly the input an agent will always max out: write high on everything and you outrank a careful writer permanently, with nothing noticing. Measuring first mattered. 45 lessons claimed high against 17 medium, but only 4 of 77 claimed it off a single record with no checkable form, so I capped those instead of refusing them. Refusing would lose real lessons to solve a five percent problem. What earns high is evidence that the mechanism transfers: it recurred across independent records, or it names something you can grep for. Certainty doesn't count, since an agent is always certain. There are two implementations on purpose, because a write-time gate never sees the files that already exist.

  11. Knowledge does not expire on a timer

    What I did, and why

    Nothing decays a lesson by age. A staleness mechanism that demoted old lessons with no recent use was proposed and rejected the same day it was designed. Rot is detected from reported harm instead: a lesson that misled a session weighs twice as heavily against as one that helped weighs for.

    Age can't tell the two kinds of old apart. A lesson bound to something the world can change — a version, a vendor, an API shape — can quietly become false. A lesson that names a mechanism doesn't expire at all, and those are the most valuable things in the corpus. Decaying by clock hits the durable ones hardest, since they're the ones that have sat longest without being touched. A stale lesson nobody retrieves costs one line in a capped budget; one that's retrieved and wrong gets reported the first time it misleads, which is also the first moment anyone could honestly know it was stale.

  12. Browsing does not count as a recall

    What I did, and why

    Agents report back what a retrieved lesson actually did, and lessons that are retrieved but never used sink while ones that changed an outcome rise. I keep queries a human types at the CLI out of those counters.

    Feedback is the only thing separating this from a pile that just gets bigger. Without it, ranking is frozen at whatever the distiller guessed on day one, and a lesson that has misled three sessions keeps outranking one that quietly saved ten. That's why the signal has to stay clean. Me poking around the corpus isn't evidence a lesson was useful, and letting it inflate the counters would corrupt the only measurement the ranking has.

  13. The site runs the retrieval engine instead of describing it

    What I did, and why

    The recall page ships the real BM25 index to the browser, and a visitor's query is ranked by the same modules the MCP server imports, copied verbatim by a generated-file step that the build checks for staleness. Nothing about the ranking is reproduced in the site's own code. It's the keyword half only: the encoder isn't loaded in the browser, and the page says so.

    Everything else on a project site is an assertion that the thing works; a console lets someone try to break it. The copy exists only because the bundler won't resolve outside the project root, and widening that root would pull the private corpus into its scope. So I generate the copy and check it for staleness, since a hand-maintained second copy of a scoring function agrees on the day it's written and drifts silently afterwards, both sides still returning plausible numbers.

  14. A new publishing surface does not inherit the old surface's filters

    What I did, and why

    Shipping the recall index to the browser would have served three client-derived lessons in full, plus their slugs as keys in the feedback map. I filtered both at the single layer that owns the publishing boundary, remapping the index's postings and recomputing its length normaliser instead of carrying them over.

    The redaction rules had been correct for every page until then, and none of that transferred to a JSON payload, because the index was built for a trusted local consumer and only became published when a new page decided to send it. The feedback map was the subtler half: it's a bag of integers, and it doesn't look like content until you notice its keys are paths and a path names a client. Found by grepping the build output for the private identifiers, not by reasoning about the code.

Stack

  • Node (ESM CLI pipeline)
  • Model Context Protocol
  • BM25 ranking
  • ONNX Runtime Web (int8 MiniLM encoder)
  • Next.js 16 (App Router)
  • React 19
  • TypeScript (strict)
  • Tailwind v4
  • Vercel
  • Model Context Protocolan MCP server exposing ten tools an agent consults per task
  • Information retrievalBM25 ranking over the distilled layer, with a fine-tuned encoder that may only reorder keyword hits
  • Enforcement designharness gates and a per-edit signature scanner that fire unread
  • Next.js 16static render over a filesystem corpus; the recall page runs the ranking engine client-side

What is honest about it

  • The corpus took months to accumulate; I built the compiler in a fast, concentrated push. The feedback loop now has four figures of recorded retrievals through it, and its first outcome-driven promotions and demotions. That's telemetry, still short of causal proof. The front page's headline stops at a refusal it can show and never says "so it can't happen again", because I haven't run the held-out replay that would earn that claim.
  • For most of one day the retrieval tools were installed and nothing was calling them: the file an adopting project loads still described the pre-rewrite design, so agents were told to grep records they're no longer allowed to read. Setup looked identical to working. It was found by using it on real work, not by reading it.
  • The gates over-fire, and I've kept the record of it: the recall gate's over-firing recurred enough times to become its own numbered series in the decision log. A gate that fires on a step not taken means every false positive taxes a session that was doing nothing wrong. I manage the tension between catching the miss and taxing the hit. I haven't solved it.
  • Scar's own README drifted from the live site's figures within a day, and later the same drift turned up in the two files every agent loads at session start, where it misleads the system instead of a reader. I now gate eleven corpus figures across those three files, plus the benchmark result, in the verify suite. Since 2026-09-28 a second check matches site sentences that a later decision made false, and checks the site's gate, tool and model rosters against the code. It only knows the claims someone listed. This page is downstream of exactly that gap: for a day after the encoder shipped, it still said Scar used no embedding model.
  • The verify suite covers retrieval quality, standing token cost, frontmatter, the confidentiality gate and the hooks. There's no test suite over the distillation stages themselves, clustering and distilling, and for a build step an agent trusts, that's the gap I'd close next.
  • The side-by-side agent demonstration on the site is a simulation; no agent is driven by it. The knowledge in it isn't: the failure shown is recorded in the corpus five times, and the retrieval beside it is a real query against the real index, run at build.
  • The encoder only runs on my machine. I haven't built the desktop app's download of the pinned model yet, so any other install gets keyword recall alone. And the 65.0% figure is a second look at the held-out set: the first served run scored 2.5 points below keyword recall, because cached vectors were decoded from the wrong byte offset. The model and how it was selected didn't change between the two runs. It's still a second look, and I'd sooner say so.
  • Scar's site lists seven models, and I'm specific about which ones do anything. One runs, the encoder. One is trained and couldn't beat shuffled labels on its last run, so it stays off and the site says it isn't better than the rule it would replace. Five are still collecting data. The site is checked so it can't call a model running unless it ships.
  • The repository stays private permanently, because the source records live in it. The distilled layer is public. That split is the product itself, and it isn't a gap in the demo I plan to close one day.

Want this kind of work on your codebase? Email me and tell me what you’re trying to build.