Wednesday, September 30, 2026

Same Engineering Habits, New AI Words

More than twenty years into my career, I reinvented separation of concerns and didn't recognize it.

I had given my AI agents job titles: Java developer, front-end developer, designer, architect, tester, DevOps engineer. The output got better, so I kept handing out hats. Then I read the research, which says a job title on its own doesn't make a model any more accurate.1 The titles weren't doing the work. Each agent had a smaller job and saw only the files that job needed. That's separation of concerns, a habit I've had since 2004, under a new name.

It wasn't the first time a pattern got ahead of its name for me. Early in my career I wrote a factory in a J2ME app before I knew the factory pattern existed, and later found my own code in a book.

my J2ME applater, I read a bookFactorypattern"Oh. That's my code."
I wrote the pattern first and learned its name later.

A year and a half into building with AI, I keep finding the same thing. The tools are new, and so are the names: context engineering, skills, agent hooks. Underneath most of them is a fundamental I already knew. And this isn't only about code. Some of my biggest lessons came from pointing agents at documents and research, and the fundamentals held there too.

The way I learn them hasn't changed either. In 2004 I read other people's code, asked whoever sat at the next desk, tried things and kept what worked. I only sometimes found out what it was called. That's how the patterns got their names in the first place: the Gang of Four wrote that their book held nothing new, only designs that already worked in more than one real system.2 It's the same loop today: try something, read what others found, notice what keeps working, give it a name, reuse it.

TryReadNoticeNameReuse2004 and 2026,same loop
The loop I learned on in 2004 is the one I still use.

The fundamentals underneath

At first the model wrote code faster than I could, so I let it. Then I read what it wrote more closely, and added guardrails. Those are CI checks by another name: automated gates every change has to pass, whoever wrote it.

On bigger projects, what the model could see decided how good its work was. That has a name now, context engineering: giving the model everything relevant, so the task is solvable at all.3 It's what I've always done when handing work to a new teammate: the background, the right documents, the constraints. With agents, I had to build that briefing as software. The project's instructions file (CLAUDE.md, AGENTS.md) is its first page, the README a new teammate reads on day one. I built retrieval over Markdown, using libraries to convert documents into text, and the hard part was making sure everything relevant actually arrived. The conversion libraries did little with images, and some of what mattered was only in the diagrams and charts. First I had to clear the clutter: meeting transcripts repeat each participant's profile photo, and emails end with the same signature. In one document library, 25,082 image references came down to 991 unique images once ingestion hashed them and skipped the repeats, and most of the 413 small ones it then dropped were avatars. I had an LLM describe the remaining 578 in plain words, and search could find what was in them.

My first knowledge graph made search worse. Its schema had effectively been designed by regular expressions: any email address became a person, and any phrase after the word "about" became a topic. The next one was small, and every part of it was there because a question needed it. With that schema in place I often didn't need retrieval at all. That's data modeling, and I had to learn it again. Now I write down the questions before I design a schema.

Around then I started turning anything I did twice into a command, a saved prompt I run by name. That's a macro, and I've always scripted anything I do twice. Producing a chapter of the reference I'm writing used to be a sequence of steps I kept in my head, and now and then I skipped one. Now it's one command that runs every step.

Commands grew into skills: a folder of instructions, and sometimes scripts, that an agent picks up when a task calls for it. Once I had more than a handful, copying them from project to project was the problem Maven and Artifactory solved for Java libraries, and the fix was the same. My skills live in a shared repository, and each project pulls the ones it needs. Agent plugins can now even declare dependencies on each other, which I wrote about as a pom.xml moment. Skills also load lazily. The agent starts with only each skill's name and a one-line description, and reads the full instructions when a task needs them.4 The skill docs call this progressive disclosure, a term borrowed from interface design. To an engineer, it's lazy loading.

Hooks were the most familiar idea of the lot. Write "never push to main" in a prompt and the model will usually listen. Put it in a hook and the question is settled before the command runs, by my code rather than the model's judgment. Hooks work in the other direction too: one can refuse to let the model stop until the tests pass. I've written pre-commit hooks for years, and these work the same way, except the change they're checking came from an agent.

Then agents learned to message each other, and I stopped being the switchboard. I automated the tedious setup (worktrees across several repos, each on its own ports) and handed the scheduling to a chief-of-staff agent. It runs sessions in parallel, lets them talk when they need to, and lines them up when one is waiting on another.

Two ideas from other people changed how I think about long runs. Boris Cherny's advice is to give the agent a goal, with /goal or /loop, and it will keep working for hours.5 A goal is a definition of done, written where the agent can check it. Andrej Karpathy's autoresearch is a loop where the agent changes something, measures it, keeps the change if the score improved, and goes again.6 He built it to improve model training, but it's hill climbing, and it fits anything you can put a number on.

There's one idea I haven't tried. People keep saying that if you encourage the model, tell it the work is good and that it can be bold, it takes on more of the job by itself. I can't tell yet whether that's real or folklore. I'll try it and report back.

NEW AI WORDSAME OLD HABIT job titles for agentsseparation of concerns guardrails on AI codeCI checks context engineeringbriefing a new teammate knowledge graphs for RAGdata modeling slash commandsmacros skill marketplacesartifact repositories progressive disclosurelazy loading agent hookspre-commit hooks /goaldefinition of done autoresearchhill climbing
New words, mostly for habits engineers already had.

The noise

The loop only works if the reading step works, and that has gotten harder. My LinkedIn and X feeds are full of sponsored posts and threads written to be clicked. "This guy built a whole SaaS in a weekend." "This paper will revolutionize everything." "Just dropped." Most of them are a screenshot and a promise, and I've learned to scroll past.

Every technology I've worked with arrived with someone promising it would change everything. The difference now is the volume and the speed. What I look for is someone who tried the thing on real work and tells you what broke.

"just dropped!!""will revolutionize everything""this guy did it in a day"your filtermost of itgoes nowherescrollpasttried it on real work, said what broke
Most of the feed is shouting. Keep what was tried on real work.

These are the people whose writing I look forward to, and I'm grateful they take the time to share it:

What's different now

Not everything has an old name. In most of the pairs above, the decision stays with me or my code: the hook blocks the push, the CI check fails the build, the schema decides what can be asked. When the model decides which tool to call next, or what to try after a step fails, I don't have an old name for that. That part is new, and I'm learning it the same way I learned everything else.

The other change is what I carry between projects. For years, a new Java service started at Spring Initializr: pick the dependencies, download a zip, build from there. Now that the model writes the code, I rarely start from a zip. I start from a short list I set up before the first prompt: guardrails, one small job per agent, a schema designed from the questions, a goal before any long run, a command for anything I do twice, skills pulled from a shared repository, and a hook for anything that must never happen. The model writes the code fresh each time. The list is what I take from one project to the next.

starter.zipthen
Recipeguardrails firstone job per agentschema from questionsgoal before long runsa command for repeatsa hook for never-againsnow
The zip used to travel between projects. Now the list does.

A list doesn't choose what to build, though. Knowing what's worth building was the hard part in 2004, and it still is. I wrote about that separately.

If you're new to this and it feels as if everyone else got a manual, they didn't. They're trying things and reading each other, the same as you. Some of them write it down, and I'm thankful they do. If there's someone you look forward to reading, I'd love to hear who.


  1. Zheng et al., "When 'A Helpful Assistant' Is Not Really Helpful", Findings of EMNLP 2024 (162 personas, no gain on factual questions); Basil et al., "Playing Pretend: Expert Personas Don't Improve Factual Accuracy", 2025. ↩

  2. Gamma, Helm, Johnson and Vlissides, Design Patterns (Addison-Wesley, 1994), Introduction: "None of the design patterns in this book describes new or unproven designs. We have included only designs that have been applied more than once in different systems." ↩

  3. Tobi Lütke proposed the term in June 2025 and Andrej Karpathy endorsed it; Simon Willison collected both. ↩

  4. Barry Zhang, Keith Lazuka and Mahesh Murag, "Equipping agents for the real world with Agent Skills", Anthropic, October 2025. ↩

  5. Claude Code's /goal keeps working until a completion condition is met; under the hood it is a Stop hook. Boris Cherny, "five tips for running Opus autonomously for hours", June 2026. ↩

  6. Andrej Karpathy, autoresearch, March 2026: the agent edits a training script, trains for five minutes, keeps the change if the result improved, and repeats. ↩

Thursday, September 3, 2026

Lessons from fine-tuning a SLM

Field report · four training runs · June 2026

Simplified Technical English didn't do it. Zinsser didn't do it. Google's style guide didn't do it. So I spent three weeks and $113 training a model to sound like the engineers I like reading. A one-line prompt to Claude beat it.

4training runs
$113total spend
11logged failures
1usable model
Openingthe bet

Forty dollars of rented H100 time, and my model finally spoke. It opened with a byline, for an article nobody had written, then repeated its own title and dropped a stray HTML comment into the middle of a paragraph.

That was run one of four.

The plan had a clean shape. A frontier model knows things, so let it carry the facts. A small model, tuned on good human technical writing, would carry the prose. Knowledge from the big one, voice from the little one.

I ran the idea past Claude before I spent a dollar. It told me, politely, that I would lose: a well-prompted frontier model would out-write a fine-tuned small one, because it is already the better writer and prompting costs nothing next to training. It was a good argument. I built the thing anyway, because I did not want to be told.

It lost. And I would spend the $113 again, which is what this post is about.

Onethe problem

The rulebooks I tried first

The technical prose these models produce is slop. Padded, hedged, structurally identical from one answer to the next, every point wearing a bold label. Once you have noticed the shape you cannot stop noticing it, and I was fighting it in every draft.

So I did not start by renting a GPU. I started with the rulebooks, because people have been codifying good technical writing for fifty years and I assumed somebody had already written the answer down.

ASD-STE100, Simplified Technical English, is the strictest: roughly nine hundred approved words, one meaning per word, short sentences. It exists so a technician reading a maintenance procedure in their third language, at two in the morning, on an aircraft, cannot misread it. At that job it is superb. It also strips the voice out on purpose, which is the entire point of it and exactly my problem.

Google's developer documentation style guide makes a large team's output consistent, which is hard and worth doing. Consistent and alive are different things.

Zinsser comes closest. But when you codify On Writing Well into something a model can follow, the half that survives is the subtractive half. Cut the adverb. Cut the qualifier. Those are rules. "Build a paragraph that carries an argument" is a craft, and it does not fit in a rubric.

Hold on to that. I am about to make the same mistake, at much greater expense.

Each of those helped, and none got me where I wanted. The output came back shorter, tighter, more correct and no more alive, because "don't sound like a machine" is a far weaker instruction than "sound like this person." You cannot write down what a good voice is. You can point at a few hundred examples of one. So point.

Twothe pipeline

Eight stages, and how each one lies to you

Tuning a twelve-billion-parameter model is a sequence, and almost nothing in the sequence fails loudly. The model rarely crashes. It hands you something slightly wrong and lets you discover it much later, usually after you have paid for the run that produced it. Every failure below happened to me.

The eight stages, and where each one lies

Left: what the stage does. Right: how it goes wrong while reporting success.

STAGE FAILS SILENTLY AS 01 Gather a corpus The one choice that is most of your result Spotless, and all reference material 02 Clean it Strip bylines, nav, cookie notices, stray HTML You remove the junk, keep the habits 03 Weight it Where your taste becomes numbers Rounding let me double a source, never halve 04 Pack and budget length Fixed-length blocks; long docs get cut The trimmed tail takes the stop token 05 Continued pretraining Soaks the model in a register, teaches no tasks Loss diverges, script still exits zero 06 Instruction tuning Trains on the answer half; shapes output most Teaches habits you never meant to teach 07 Merge and shrink Fold in the adapter, quantise: 24 GB to about 7 Seven tooling failures, none about the model 08 Judge it honestly Generate on unseen prompts and read the output You measure loss, because loss is easy to plot The stages interact: a weighting choice in 03 changes the sound in 06.
Every setting is a hypothesis, and you learn which ones were wrong all at once, at the end.
Threethe corpus

Your model is a mirror of your corpus

Run one finished cleanly and produced garbage. It wrote author bylines, repeated its own title, left HTML comments sitting in the prose. I had scraped a few hundred sources and then taught the model, carefully and at some expense, to imitate the scrape.

The audit was humbling. My data preparation code was prepending a markdown byline to 89% of the training documents. Nearly a fifth of my instruction examples ran longer than the context window, so the trainer cut their tails off and took the stop token with them. That is why the model never knew when to stop talking.

The subtler failure showed up after I had fixed all of that. The corpus passed every cleanliness check I could write and was still wrong, because four-fifths of it by word count was dry specification text. I had assembled it for coverage when the model's entire job was register. Rewriting the weighting took readable human voice from 16.5% of the corpus to 59.7%, and made it smaller doing it: 11,546 packed blocks down to 7,034.

Then someone asked me a plain question mid-run: will it stop answering with links? I went and looked instead of guessing, and 84% of my instruction answers contained inline markdown links, six apiece at the median. I killed the run there, two dollars in, stripped the URLs and started over. Letting it finish would have cost $9 and produced something I would have thrown away. One plain question was worth more than every automated check I had written.

Fourthe failure

The run that went perfectly and produced a wreck

The second attempt trained to completion, published its result and reported success. Final evaluation loss of 7.37, worse than where an untrained model starts. Training dropped the way it should, bottomed out around 1.48 by step 30, then blew up and stayed up for the remaining seven hundred steps. A learning rate one notch too high. Nothing errored. It spent the money and handed me the wreck with a clean exit code.

Three separate safety nets should have caught this. All three were looking somewhere else.

Where the guards were looking

Run v2b, lr 2e-4. Loss values are from the run's history; guard positions are from the config.

8 6 4 0 0 400 722 training step loss smoke test stopped here step 40 · $1 spent · all healthy divergence starts ~step 100 a kill switch here aborts at $3 both saved checkpoints already broken run completes · $19 · publishes the wreck
The smoke test ran forty steps and saw only the healthy opening descent. The quality gate that generates real text crashed on a tokenizer quirk, and my own rule said to treat a crashed check as infrastructure and carry on. The trainer's keep-the-best-checkpoint setting had nothing good to keep.

A smoke test that ends before your failure mode starts only reassures you. So I stopped and rebuilt the harness before spending another dollar: a kill switch inside the training loop, milestones with the billing switched off between them, and uploads verified by re-reading the destination rather than trusting an exit code.

Fivethe bill

What it cost, run by run

Three of the four runs produced nothing I kept. That ratio is roughly what learning something looks like.

Four runs, one usable model

Rented H100 time at $3.29/hour, June 2026.

v1poisoned corpus
$40
v2 first tryuploads failed silently
$21
v2 second trydiverged, stopped
$19
v3trained clean, published
$33.50
The $21 run still stings: it trained correctly, then every upload failed quietly because the command was written to ignore its own errors, and the machine destroyed itself on schedule with the only copy on it. An H100 also sat idle for 1.6 hours while I made one decision, which cost $5.30 for nothing.
Sixthe verdict

The test that actually mattered

The third run trained clean and measurably cut the machine-writing tics: 0.91 on my LLM-speak metric against a stock Gemma's 1.33, winning thirteen of twenty held-out briefs. I had built a voice tool rather than a smarter model, which is what I set out to build.

Then I gave the same notes and the same plain instruction to three writers: Claude Sonnet, a stock Gemma 4 26B, and mine. Then I read all three.

Claude Sonnet

Frontier · plain prompt · 972 words

"When you stuff a complex task into a single prompt, you're asking the model to do too many things at once. It will drop instructions, mix up concerns that should stay separate, and produce output that's difficult to debug."

Every word inside a paragraph. Not one bullet in the whole chapter.

Gemma 4 26B

Stock · plain prompt · 874 words

"…you gain three massive advantages:

* Reliability: Each call is simpler…
* Observability: You can inspect and test…
* Granular Control: You aren't locked in."

Opens with good paragraphs, then reaches for a bold label every other line. More than half its words live inside bullets.

My fine-tune

Tuned for voice · 422 words

"Your LLM prompt is too big and unwieldy for the task."

"Your prompt has a lot of conflicting instructions."

"Your prompt's output is hard to parse."

No filler, no hedging, and no paragraphs either. I asked for a chapter and got a list of sentences standing on their own.

Average length of a paragraph, in words

Measured on the three chapter files. Headings excluded.

Sonnet
44.9
Gemma 26B
40.1
My fine-tune
10.8
At eleven words, those are sentences with line breaks between them. My model had learned to strip the padding out of a sentence and had never learned to build anything out of sentences. Sonnet put 943 words in paragraphs and zero in bullets; Gemma split 361 against 487.

So the model I trained specifically for voice lost on voice, to a model I simply asked nicely.

Seventhe reason

Why it lost

I had taught it the easy half of writing well. It dropped the hedges, the throat-clearing, the "it's important to note." Those are words, and words are cheap to count, so words are what my scoring rewarded.

Structure is the hard half, and structure is the deeper tell. Bulleting everything is what makes machine writing feel like machine writing, more than any individual word does, and my corpus was still full of documents that bullet everything. It taught the structure right back in while I was congratulating myself about vocabulary. I optimized the symptom I could count and missed the one I could not. Which is the same mistake the rulebooks make, run again at $113 a go. Subtraction fits inside a rule. Construction needs judgment.

A frontier model, asked nicely, wrote better prose than three weeks of training did.
Eightwhat to take

What I owe the people who do this for a living

Every mistake here is an ordinary, known thing to someone who trains models for a living. A practitioner would have started the learning rate lower and clipped the gradients without thinking about it. A clean dataset can still be the wrong dataset, which is the first thing anybody who has built one will tell you. I paid $113 to learn a page of things a good team already knows in their hands.

And every choice in the recipe is a hypothesis. I made reasonable guesses and ran four configurations. The people who are good at this run hundreds, changing one variable at a time, and most of those runs teach them nothing except that a door is closed. Patient, largely unrewarded, mostly negative-result work. That is the actual job.

Almost none of what I used did I pay for. The loaders, the trainers, the quantiser, the hub the model sits on: built and given away by people who did not have to. So were the essays my corpus is made of, a few hundred engineers who wrote carefully in public for no particular reward. Whatever voice my model has, it is theirs. I borrowed it.

What $113 bought me

The eight rules I would hand to anyone starting out, in the order they cost me money.

THE RULE WHAT IT COST TO LEARN Have a specific reason. A rulebook only buys you a floor. Read your data. Audit for what the model will copy. Make the smoke test reach the failure. Put the kill switch inside the loop. Verify the upload landed. Answer plain questions with data. Privacy, cost, latency, offline, or a narrow repeating task. Style guides say what to remove, never what to build. Falling loss means it predicts your text, garbage included. Tone and structure count as contamination, same as markup. Length it to where runs historically break. A divergence check aborts early; a finished run publishes. An exit code of zero says nothing about the bytes. "Will it still do X?" is a validation prompt. three weeks the whole idea $40 · run v1 $40 · run v1 $19 · run v2b $16 of that $19 $21 · run v2a saved $9 for $2
The last row is the only one in the black. Going to look cost two dollars and saved nine.

If you want this voice today, prompt a frontier model. One line of instruction got me cleaner prose than three weeks of training did. If you want a fine-tune that genuinely wins on voice, I think it is findable and I did not find it: a corpus weighted hard toward long-form essayists, an explicit penalty against list structure, and a judge that scores the shape of the prose instead of which words show up in it. Whether that works, I don't know.

Claude told me I would lose, and Claude was right. Being told cost me nothing and taught me nothing. Losing cost me $113 and three weeks, and I now understand with my hands, rather than in the abstract, why training is hard and why the people who are good at it are worth listening to closely. The models will keep getting better at telling us things. We will keep needing to go and find out.

The model is public if you want to poke at it, and the full list of what went wrong is much longer than this post. If you have run this experiment and found the corpus that wins on voice, I would like to hear about it.

The tuned model, adapter and merged weights: huggingface.co/abdwiv/gemma4-12b-tech-writer
Base model Gemma 4 12B · two-pass LoRA · trained on rented H100 · June 2026 · every figure here measured from the run logs and the output files.

Sunday, June 14, 2026

The Taste You Can't Outsource

It was late, and I was doing the kind of work that never makes it into a demo: adding guardrails to my Claude Code setup. While I was in there I pulled in SkillSpector, NVIDIA's security scanner for AI agent skills. It checks a skill for malicious patterns before you let it near your machine. The docs were stale and a couple of things were broken, so I did what I do now. I asked Claude what else was missing.

It came back with two recommendations. The second one stopped me cold.

Remove the call to OSV. Add an offline mode that doesn't reach the internet.

Wait, what is OSV, and why does it even need to connect? OSV is the Open Source Vulnerabilities database, a free public service (osv.dev) that maps known security flaws to specific package versions. When SkillSpector spots a dependency, it asks OSV one question: is this exact version known to be vulnerable? That single call is how the scanner knows what "bad" looks like today, instead of whatever happened to be true the day the code was written.

So for a scanner whose whole job is to catch known-bad code, the call to OSV isn't a feature. It's the part that does the looking. A mode that skips it isn't a leaner tool. It's still a cheese burger - with cheese and burger, just without the beef.

And the suggestion wasn't wrong, exactly. I pushed on it, and Claude made a coherent case: air-gapped CI, no network egress, faster runs. Every one of those is real. In a different tool it would be good advice. The model wasn't hallucinating. It was reasoning. It was just reasoning about everything except the one thing that made the tool worth building.

Build anything. In a day.

We're deep in the season of the grand claim. AI will replace engineers. You can ship a feature-complete product in an afternoon. There's a skill that turns an agent into your chief of staff, and a thread every week where someone stands up a whole app over a weekend and a thousand replies ask for the prompt.

I want to be generous, because the capability underneath is astonishing and I use it daily. But I think we're mistaking a capability demo for a product. Those builds are samples. They show what the clay can do. They are not the same as knowing what to make from it. And a product was never "what the model can build." It's what you wanted it to build. Different sentence.

The taste it can't have

The model has taste. Ask Claude to make something nice and it will. What it doesn't have, and I'd argue can't, is taste specific to you: to the single reason this thing exists and not some adjacent thing that would also be defensible.

That reason isn't in the code. It's in the point. And the point lives in your head, not the repository. So the model optimizes what it can see, like "faster" or "more flexible" or "offline," and quietly trades away the thing it can't: this is a security tool, and a security tool that doesn't check is worse than no tool, because it returns green without looking.

Here's the part that got under my skin. I'm proud of my Claude setup. It knows my preferences, my level, the work I do; it doesn't hand me the vanilla answer. By any measure it's well-grounded in me. And it still told me to unscrew the sensor. Which means this isn't a prompting problem you tune your way out of.

Knowing what to build is the job now

So who does well here? Not the fastest prompter. The person who can put on the product hat and keep the engineering skill to get it done, and knows which is which.

Knowing SkillSpector must call OSV is product knowledge. It's a judgment about what would, and wouldn't, bring value, and it's exactly the judgment the model skipped. The engineering question is what you reach for after: SkillSpector already falls back gracefully when OSV is unreachable, and that's the careful version of "offline." Deciding the database is optional is not the same as handling the day it's down. One is a product decision. The other is engineering.

And the engineering is the part I'm actually building right now. A scanner only protects you if you remember to run it, so I'm putting it in front of the door: a guardrail that checks a skill before it ever installs, and hands back a clean allow, ask, or deny. To the agent reaching for a new skill, and to the CI pipeline doing the same thing on a human's behalf. The OSV call stays non-negotiable; what gets easier is everything around it. Telling those two apart, the line you must never cross and the capability you can keep extending, is becoming the real skill. More on the build another time.

What I'm not saying

This isn't an "AI is overhyped" piece; I don't believe that, and the story doesn't support it. The model found real bugs in that library, fixed the stale docs in seconds, and its other recommendation was good. I shipped it. On the how, it's a genuine force multiplier.

But the harder a thing is to write down, and the reason a tool exists is almost impossible to write down, the longer it stays ours. So I screwed the sensor back in, kept the OSV call, and left the dangerous advice on the floor. At the end of a late night of plumbing I didn't feel threatened. I felt useful. The model could build almost any version of that tool I asked for. It just needed me to know which one was worth building.


Ideated and dictated by me, written by Claude

Thursday, June 4, 2026

A brief history of plugin.json

A Brief History of plugin.json (Claude Code)

A Brief History of plugin.json

The evolution of Claude Code's plugin manifest system into a full-fledged dependency management engine.

1. The Era of Fragmentation (Late 2025)

When Anthropic first stabilized the Claude Code CLI and the Model Context Protocol (MCP), extendability was highly fragmented.

  • MCP Servers handled external tools (APIs, databases, filesystems).
  • Settings files (.claude/settings.json) handled rules and configurations.
  • Skills (like custom prompts) lived as standalone instruction documents.
  • No unified concept of a "packaged extension" existed; workflows had to be manually wired together via scripts.

2. The Birth of the Manifest: .claude-plugin/plugin.json (Early 2026)

Anthropic introduced the Claude Code Plugin Architecture to unify components into a singular standalone structure.

  • Plugins bundled skills, custom sub-agents, hooks, and local .mcp.json tool declarations.
  • Introduced plugin.json as a passive metadata descriptor handling basic identity and simple version strings.
  • Versioning remained loose, resolving primarily by binding a project path to whatever git SHA happened to be HEAD at runtime.
{
  "name": "deploy-kit",
  "description": "Handles infrastructure provisioning and AWS EKS hooks",
  "version": "1.0.0"
}

3. The Broken Cache & Chaos Crisis (Spring 2026)

As enterprise adoption scaled, major operational and environmental cracks emerged in production environments.

  • Non-Deterministic Environments: Shifting git SHAs caused different team members to resolve different variations of the same plugin.
  • Cache Nightmares: Broken builds cached locally (~/.claude/plugins/cache/) persisted indefinitely due to a lack of native self-healing mechanisms.
  • Silent Failures: Master plugins relying on utility plugins had no mechanism to declare relationships, leading to fragile manual setup guides.

4. The Modern Era: Pure Parity & Version Constraints (June 2026)

Anthropic rolled out native Plugin Dependency Resolution (v2.1.143+), transforming plugin.json into an active package manager specification heavily inspired by Node's package.json.

  • Graph Enforcement: The CLI actively blocks disabling a plugin if active core systems rely on it, tracking the entire transitive dependency graph.
  • Cross-Marketplace Guardrails: Auto-installing dependencies across disparate marketplaces is locked down by default to prevent supply-chain attacks.
  • Strict Release Tagging: Enforced tag pushing (claude plugin tag --push) matches git tags directly to the manifest version.
{
  "name": "deploy-kit",
  "version": "3.1.0",
  "dependencies": [
    "audit-logger",
    {
      "name": "secrets-vault",
      "version": "~2.1.0"
    }
  ]
}

This structural shift marks the transition of AI tools away from loose "prompt scripts" and into production-grade, tightly governed software engineering components.

All content above provided by courtesy of Gemini. No claims mine.

Claude Code Shipped a Dependency Manager, and I Think It Is a New Frontier

Claude shipped a plugin dependency manager

Claude Code recently shipped dependency management for plugins, and as someone working on making collaboration better, I feel very excited about it. It means a skill can now be built on top of another skill. It means a platform team can publish a foundation and a hundred other teams can build on it without copying a single file. An entire layer of reuse that simply was not possible a few weeks ago is now sitting in the changelog, described in one modest sentence. I read it twice, and then I sat back, because I have seen this exact moment before in other corners of our industry, and I know what comes after it. It is a new frontier. Think about the introduction of pom.xml, or of package.json.

A quick word on how this post came together. I worked through most of this thinking out loud with Claude itself, tracing the feature from the docs into the actual GitHub issues, arguing about what counts as a dependency manager and what does not. Some of the sharpest framing here came out of that back and forth. I mention it because the irony is not lost on me that I used the tool to understand the tool.

Why those two files are the right comparison

The reason I reach for Maven and npm is that I have leaned on both, and the contrast between them taught me what dependency management actually buys.

In Maven, a dependency is an intent rather than a list of files.

<dependency>
  <groupId>org.springframework.boot</groupId>
  <artifactId>spring-boot-starter-web</artifactId>
  <version>3.2.0</version>
</dependency>

I ask for one artifact and Maven brings in dozens, because Spring Boot declares its own dependencies, and those declare theirs, all the way down. The transitive set arrives without my touching it. The project owns a pom.xml, and when the build runs, Maven reads that file and provisions the world it describes. The project is the thing that sets resolution in motion.

npm expresses the same idea, with one addition I have come to value.

{
  "dependencies": {
    "express": "^4.18.0",
    "pino": "~9.0.0"
  }
}

Those carets and tildes are quietly profound. They are a contract about change. The caret welcomes compatible updates but refuses to cross a major boundary in silence. When I install, npm intersects every package's constraint, settles on versions that satisfy all of them at once, and writes the exact resolved set into a lockfile so the next machine reproduces it down to the patch. The constraint expresses tolerance for change, and the lockfile expresses intolerance for surprise, and a mature ecosystem needs both.

Both files, underneath the syntax, do the same four things. They let a consumer declare what it needs, version what it needs, resolve the transitive graph so one install pulls in everything below it, and reproduce that resolution elsewhere. Declare, version, resolve, reproduce. That quartet is what turned two configuration files into the foundation of entire economies, and it is the lens I want to hold Claude Code up against.

What actually shipped

https://github.com/anthropics/claude-code/issues/64457

A plugin now declares its dependencies in its manifest, in a form anyone who has opened a package.json will recognize.

{
  "name": "platform-observability",
  "version": "2.3.0",
  "description": "Shared observability conventions and tooling",
  "dependencies": [
    "audit-logging",
    { "name": "metrics-core", "version": "~2.1.0" }
  ]
}

There are two forms, and the choice between them is the same lesson npm teaches. A bare name accepts whatever the marketplace currently offers. An object with a version field pins a semantic version range, and the highest release that satisfies it is selected.

Now a higher-level plugin builds on that foundation.

{
  "name": "team-incident-toolkit",
  "version": "1.0.0",
  "dependencies": [
    { "name": "platform-observability", "version": "^2.0.0" }
  ]
}

Someone installs the leaf.

claude plugin install team-incident-toolkit@acme-marketplace

The runtime reads the manifest, sees the declared dependency, and pulls in platform-observability automatically, and with it the things that plugin in turn depends on. One command, the entire closure. The top of the tree arrives with the tree attached. Versions are anchored to git tags on the marketplace repository, and when two installed plugins constrain the same dependency, the runtime intersects their ranges and chooses the highest version satisfying both, raising a clear error when no version can satisfy everyone rather than quietly loading something that will misbehave later. That last behavior is the one I respect most, because the failures that hurt are never the loud ones.

Three of the four properties are unmistakably here. Declaration, versioning with real semantic ranges, and transitive resolution with honest conflict handling. The fourth, reproducibility, is where I have to temper the enthusiasm, because there is no lockfile yet. Resolution happens against live tags at the moment of install, so reproducibility rests on tight pinning by convention rather than on a committed artifact that guarantees every engineer and every build agent lands on an identical set. For anyone who has come to treat a lockfile as the thing that makes a teammate's checkout match their own, its absence is the conspicuous gap, and I expect it to be among the first things to close.

It is worth noting that this did not arrive fully formed, and the community is actively shaping it. The original request lived as issue #9444, asking for exactly this, plugins that declare dependencies on other plugins, with a shared library plugin underneath. Reading that thread is a small lesson in how these features mature in the open. And the rough edges are being found the same way. The version resolution itself was reported broken in issue #64457 and, as I write this, is only partially fixed, with local folder marketplaces still misbehaving, tracked in issue #65337. If you build on this today, build on a remote marketplace and verify your version pins resolve before you trust them.

So far this is a celebration with one footnote. Now I want to turn to the two things that keep me from calling this finished, because a frontier is exciting precisely because it is not yet settled.

It is Claude's, not a standard, and I hope that changes

The first thing to sit with is that this is Claude Code's dependency model, and Claude Code's alone. The manifest format, the tag convention, the resolution behavior, all of it lives inside one vendor's tooling. That is not a criticism. It is how every one of these stories begins. npm was a JavaScript thing before it was a movement. Maven was a Java thing. A dependency model almost always arrives as one ecosystem's local invention and only later, if it earns it, hardens into something the wider world agrees on.

But I find myself wishing for the standard already, because the underlying artifact is more portable than the plumbing around it. The skill itself, the folder with its instructions, is largely tool-agnostic. What differs is the wrapper. A different agent tool expresses the same dependency idea in a different manifest with a different resolver, and a team that lives across more than one tool ends up maintaining two disciplines for one set of skills. We have watched a parallel effort try to standardize the skill format itself. I would love to see the same energy reach the dependency layer, so that "this plugin depends on that plugin, in this version range" means the same thing no matter whose runtime reads it. We already have a standard for how agents talk to tools, and the field is converging on one for how agents talk to each other. A shared way to express how agent capabilities depend on one another feels like the natural third piece, and I hope someone takes it up. Until then, what we have is a very good vendor implementation, and a very good vendor implementation is exactly the seed a standard grows from.

There is no lifecycle, so the trigger is still a person

The second thing is subtler, and it is the one that took me a moment to articulate. Go back to pom.xml for a second. The reason that file is powerful is not only that it declares dependencies. It is that a lifecycle reads it. I run a build, and the build looks at the project, sees what it needs, and provisions it. The project itself is the trigger.

Claude Code has the declarations and the resolution, but not that lifecycle. The dependencies belong to the plugin, and resolution fires when a person installs a plugin, not when a repository is opened. A project can describe the plugins it wants, but nothing yet reads that description on open and provisions it the way a build reads a pom.xml. The engine is solid. What is missing is the project acting as the trigger.

In practice this means onboarding still begins with a deliberate human act. A new teammate clones a repository, and the right capabilities do not simply appear because the repository asked for them. Someone has to issue the first install. The dependency graph beneath that first install is fully automatic and genuinely impressive. The initiating step is not. It is a small manual seam stitched on top of a hard problem that has already been solved, and that ratio is exactly why I am optimistic rather than impatient. The expensive machinery is built. What remains is a convenience, and conveniences tend to follow quickly once the hard part is done.

Why I bothered to write this down

A changelog entry says plugins can declare dependencies. I read the same line and saw the thing I have learned to recognize.

The reason npm and Maven grew into economies was never the syntax. It was that dependency resolution is what turns a shared library into something a hundred teams can build on without copying it. Once that primitive exists, a platform team can own a foundation, publish it, version it, and let others compose on top while pinning the compatibility they have actually tested. That is now possible for agent capabilities, and it was not a month ago. A platform group can own a shared plugin, conventions and wiring and a few well-tested skills, and another team can depend on it in a single line and receive updates inside a range it trusts.

The question for those of us who think in systems is no longer how to share a skill, which has quietly become an ordinary solved problem. It is how to architect a layered set of capabilities the same way we architect a layered set of libraries. Stable foundations underneath, faster-moving work at the edges, version contracts between teams who do not sit together. These are old disciplines, and they have found a new substrate.

I will close with the honesty this deserves, because the ground is still moving as I write. The behavior I have described landed recently. The missing lockfile, the vendor-specific format, the install-time trigger that has not yet become a project-time one, all of it is under active iteration. What I have written is a snapshot. The durable point is not today's exact feature set but the direction, which is that agent tooling is compressing into months the same evolution that language ecosystems took years to work through. We are watching a pom.xml moment happen in real time. I do not know exactly what it grows into, and I am genuinely excited to find out.

And FWIW, I am excited about the analysis and havent tried it yet.

If this resonated, I would love to hear your perspective, especially if you are an engineer or engineering leader thinking about how shared capabilities should be built and distributed in an AI-first world.

Monday, March 16, 2026

Thunderstorms and Sunshine | A Principal Software Engineer's perspective

By a Principal Engineer, March 2026

It's a mid-March morning. Sunny, a little cloudy, with a pleasant breeze in the air. I sit here with my coffee, thinking about software engineering: where it's been, where it's going, and what AI really means for those of us who have given our lives to this craft. It's a calm morning. But I know what's coming. By afternoon, severe thunderstorms. Tornadoes on watch. Schools closing early.

Tomorrow, though, will be beautiful again.

That's where we are with AI and software engineering. And I think it's worth talking about honestly. Not with hype, not with fear, but with the perspective of someone who has been in this industry long enough to have seen a few storms before.

Where It Started

I was in 8th grade when I fell in love with software engineering. Not through a class, not through a mentor. Through a GW-BASIC program I found printed somewhere, typed up by hand, and ran on a DOS machine.

On the screen, an apple tree appeared. Apples grew, circled, disappeared, grew back. It was visually enigmatic. Beautiful. And the thing that hit me wasn't "I ran a program." It was: I made something beautiful that wouldn't have existed otherwise.

That feeling is one I've been chasing ever since.

I grew up, got my education, started engineering professionally. As a junior, you learn a lot and influence little. You're a small piece of something enormous, and that's humbling in a good way. But every now and then, something happens that cements why you're here. For me, one of those moments was tracking down a bug that had been crashing production servers for months. Objective-C code. An invalid memset, hiding in plain sight. It wasn't easy to find. It took patience, stubbornness, and a refusal to accept "we don't know why it's crashing." When I finally found it and fixed it, the joy was extraordinary. Not because it was glamorous. Because it was hard, and it mattered.

Those early moments shaped how I see software engineering. It is never easy. It is never one-shot. But there is absolute joy in doing things that would otherwise be very difficult to do.

The Drift, and the Return

Over the years, I moved deeper into architecture, design, and management. My teams wrote the code. I shaped the thinking, made the calls, unblocked the hard problems. I'd still jump in when something needed to move faster than anyone else could move it. That instinct never leaves you. But the hands-on building became rarer.

Then, in February 2025, vibe coding arrived.

I'd experimented with AI-assisted coding before. You'd describe a problem, get some code back, it was interesting but inefficient. More of a curiosity than a capability. Vibe coding changed that. For the first time, I could sit down with an idea and build, really build, with AI as a genuine collaborator.

I started with a side project. And within hours, I felt it again. The same joy. The apple tree from 8th grade. The feeling of making something that wouldn't have existed without me.

Here's what I think vibe coding actually unlocked: it removed a specific kind of friction that had accumulated over years. Not the intellectual friction, which is the fun part. The volume friction. The syntax differences between languages. The boilerplate. The fact that you can see exactly what needs to exist, you know it down to your bones, but materializing it takes days. AI collapsed that gap. And in doing so, it gave back the builder's joy to people like me who had drifted away from the code.

We don't write code because writing code is the end goal. We write code to build things. To make something real out of an idea. AI made that more accessible, not less meaningful.

Speed Goes Up. Judgment Matters More.

Let's be clear about something. AI accelerating code output does not mean engineering judgment matters less. It means it matters more.

Software engineering has always been both science and art. Over decades, our industry has accumulated hard-won principles: DRY, SOLID, 12-factor. Not arbitrary rules, but distilled lessons from countless projects that went sideways. A senior engineer doesn't always recite these principles by name. But they feel them. They look at a piece of code and know, almost instinctively, whether it's going to cause pain in six months.

That instinct doesn't transfer to AI. Not yet.

Here's a real example. Recently, I was reviewing AI-generated code that needed to process log data. The AI was stuck trying to read those logs top-to-bottom. It kept grinding away at the problem that way because that's the obvious path. But anyone who looked at how that log was actually structured would immediately know: read it bottom-to-top. That's it. Problem solved. The AI couldn't see that because it was constrained by its framing of the problem. I wasn't.

Another example: a developer on my team was running into issues where an AI was failing to generate detailed outputs for a batch of work items. His instinct was to increase the context window, to just throw more at it. Classic junior mistake, honestly. The right move was the opposite: break the problem into smaller chunks, give AI manageable pieces, and work through them sequentially. The same judgment I'd give a junior engineer, I now give to AI-assisted workflows. The nature of the guidance hasn't changed. The recipient has.

AI is an amplifier. If your thinking is good, it makes you significantly more productive. If your thinking is flawed, it produces flawed code faster. The engineering judgment, the taste, the architecture, the "wait, why are we even approaching it this way" instinct: that is not being replaced. It is being put to work more than ever.

The Factory Problem

Last week I saw a post about building a software engineering factory: end-to-end automated delivery of software. The ideas were genuinely interesting and I respect the ambition. But I keep coming back to something.

I graduated as a civil engineer before becoming a software engineer. In civil engineering, once the design is fixed, it's fixed. You don't change the foundations of a bridge mid-construction. The beauty and the complexity of software engineering both arise from the same source: it is fluid.

Scope creep exists because it can exist. Agile and SAFe came into being because until people actually see software running, they don't fully know what they want. We are great as humans at imagining abstract ideas. We are not great at knowing exactly how we want them realized until we can see and touch them. Two architects in a room will have an animated, sometimes heated discussion about design tradeoffs. That's not a bug. That's the process.

A factory model will deliver something standardized. That's valuable, maybe for 40 or 60% of use cases. But the cases that stand out, the products that resonate, the software that people love, those come from someone having a point of view. A taste. An opinion about what this specific thing should feel and do and be.

AI can build. But AI cannot want. It cannot tell you which vision is worth building. It cannot feel whether something will resonate with your specific audience in your specific context. That judgment, the product mindset, the vision, the "this is what I want it to be and here's why," is irreducibly human. You wouldn't delegate that to a developer without product sense. You wouldn't delegate it to a PM who doesn't understand the users. And you shouldn't delegate it to AI either.

The engineers and leaders who will thrive in this era are the ones who develop that product sensibility alongside their technical depth. Not one or the other. Both.

The Question I Don't Have an Answer To

Here's the thing that keeps me up at night. I want to be honest that I don't have a clean answer.

Taste is earned. You don't graduate as an architect. Nobody hands you the instinct for good system design. You build it over years: through production bugs, through painful refactors, through the memset hunts and the log-reading optimizations and the countless small decisions that add up to something called experience.

If AI absorbs more and more of that work, the debugging, the boilerplate, the small architectural decisions, where does the next generation of principal engineers come from? How do you develop taste without the struggle that forges it?

I don't know. I think there will be mistakes. Some will be caught by the architects and principal engineers who still have the eye for it. Some will make it to production. Companies I deeply respect have already shipped AI-assisted code that caused real revenue impact. That will keep happening for a while.

I believe it will get better. My rough mental model is that we are at step 3 of this evolution. Step 0 was humans writing everything. Step 1 was AI assistance. Step 2 was agentic and vibe coding. Step 3, where we are now, is AI that is starting to develop a kind of flavor. A taste. It doesn't always get it right, but sometimes it does, and you can feel the difference. In five years, I think AI will write reliably good code. In ten years, production-grade code without heavy supervision. But we are not there yet. And in the gap, human judgment is not optional.

Tomorrow Will Be Beautiful

This morning started with sunshine and a pleasant breeze. By afternoon, severe thunderstorms are coming. Tornadoes. Schools are closing early.

AI is going to speed things up. It is going to cause chaos. Teams will shrink. Roles will shift. Some of what we've taken for granted about how software gets built will be upended. That disruption is real, and pretending otherwise helps nobody.

But here's what I'm confident about: software engineering as a profession is going to stay. The number of people will change. The skills that matter will shift. But the need for humans who have taste, who have a vision, who know what good looks like, who can tell the difference between code that will hold up and code that will crumble, that need is not going away. If anything, as AI makes building cheaper and faster, the premium on knowing what to build and why goes up, not down.

The thunderstorms are coming. I'm not going to pretend they won't be disruptive.

But tomorrow is going to be a beautiful day.


If this resonated, I'd love to hear your perspective, especially if you're an engineer or engineering leader navigating these same questions.