Showing posts with label #ClaudeCode. Show all posts
Showing posts with label #ClaudeCode. Show all posts

Wednesday, September 30, 2026

Same Engineering Habits, New AI Words

More than twenty years into my career, I reinvented separation of concerns and didn't recognize it.

I had given my AI agents job titles: Java developer, front-end developer, designer, architect, tester, DevOps engineer. The output got better, so I kept handing out hats. Then I read the research, which says a job title on its own doesn't make a model any more accurate.1 The titles weren't doing the work. Each agent had a smaller job and saw only the files that job needed. That's separation of concerns, a habit I've had since 2004, under a new name.

It wasn't the first time a pattern got ahead of its name for me. Early in my career I wrote a factory in a J2ME app before I knew the factory pattern existed, and later found my own code in a book.

my J2ME applater, I read a bookFactorypattern"Oh. That's my code."
I wrote the pattern first and learned its name later.

A year and a half into building with AI, I keep finding the same thing. The tools are new, and so are the names: context engineering, skills, agent hooks. Underneath most of them is a fundamental I already knew. And this isn't only about code. Some of my biggest lessons came from pointing agents at documents and research, and the fundamentals held there too.

The way I learn them hasn't changed either. In 2004 I read other people's code, asked whoever sat at the next desk, tried things and kept what worked. I only sometimes found out what it was called. That's how the patterns got their names in the first place: the Gang of Four wrote that their book held nothing new, only designs that already worked in more than one real system.2 It's the same loop today: try something, read what others found, notice what keeps working, give it a name, reuse it.

TryReadNoticeNameReuse2004 and 2026,same loop
The loop I learned on in 2004 is the one I still use.

The fundamentals underneath

At first the model wrote code faster than I could, so I let it. Then I read what it wrote more closely, and added guardrails. Those are CI checks by another name: automated gates every change has to pass, whoever wrote it.

On bigger projects, what the model could see decided how good its work was. That has a name now, context engineering: giving the model everything relevant, so the task is solvable at all.3 It's what I've always done when handing work to a new teammate: the background, the right documents, the constraints. With agents, I had to build that briefing as software. The project's instructions file (CLAUDE.md, AGENTS.md) is its first page, the README a new teammate reads on day one. I built retrieval over Markdown, using libraries to convert documents into text, and the hard part was making sure everything relevant actually arrived. The conversion libraries did little with images, and some of what mattered was only in the diagrams and charts. First I had to clear the clutter: meeting transcripts repeat each participant's profile photo, and emails end with the same signature. In one document library, 25,082 image references came down to 991 unique images once ingestion hashed them and skipped the repeats, and most of the 413 small ones it then dropped were avatars. I had an LLM describe the remaining 578 in plain words, and search could find what was in them.

My first knowledge graph made search worse. Its schema had effectively been designed by regular expressions: any email address became a person, and any phrase after the word "about" became a topic. The next one was small, and every part of it was there because a question needed it. With that schema in place I often didn't need retrieval at all. That's data modeling, and I had to learn it again. Now I write down the questions before I design a schema.

Around then I started turning anything I did twice into a command, a saved prompt I run by name. That's a macro, and I've always scripted anything I do twice. Producing a chapter of the reference I'm writing used to be a sequence of steps I kept in my head, and now and then I skipped one. Now it's one command that runs every step.

Commands grew into skills: a folder of instructions, and sometimes scripts, that an agent picks up when a task calls for it. Once I had more than a handful, copying them from project to project was the problem Maven and Artifactory solved for Java libraries, and the fix was the same. My skills live in a shared repository, and each project pulls the ones it needs. Agent plugins can now even declare dependencies on each other, which I wrote about as a pom.xml moment. Skills also load lazily. The agent starts with only each skill's name and a one-line description, and reads the full instructions when a task needs them.4 The skill docs call this progressive disclosure, a term borrowed from interface design. To an engineer, it's lazy loading.

Hooks were the most familiar idea of the lot. Write "never push to main" in a prompt and the model will usually listen. Put it in a hook and the question is settled before the command runs, by my code rather than the model's judgment. Hooks work in the other direction too: one can refuse to let the model stop until the tests pass. I've written pre-commit hooks for years, and these work the same way, except the change they're checking came from an agent.

Then agents learned to message each other, and I stopped being the switchboard. I automated the tedious setup (worktrees across several repos, each on its own ports) and handed the scheduling to a chief-of-staff agent. It runs sessions in parallel, lets them talk when they need to, and lines them up when one is waiting on another.

Two ideas from other people changed how I think about long runs. Boris Cherny's advice is to give the agent a goal, with /goal or /loop, and it will keep working for hours.5 A goal is a definition of done, written where the agent can check it. Andrej Karpathy's autoresearch is a loop where the agent changes something, measures it, keeps the change if the score improved, and goes again.6 He built it to improve model training, but it's hill climbing, and it fits anything you can put a number on.

There's one idea I haven't tried. People keep saying that if you encourage the model, tell it the work is good and that it can be bold, it takes on more of the job by itself. I can't tell yet whether that's real or folklore. I'll try it and report back.

NEW AI WORDSAME OLD HABIT job titles for agentsseparation of concerns guardrails on AI codeCI checks context engineeringbriefing a new teammate knowledge graphs for RAGdata modeling slash commandsmacros skill marketplacesartifact repositories progressive disclosurelazy loading agent hookspre-commit hooks /goaldefinition of done autoresearchhill climbing
New words, mostly for habits engineers already had.

The noise

The loop only works if the reading step works, and that has gotten harder. My LinkedIn and X feeds are full of sponsored posts and threads written to be clicked. "This guy built a whole SaaS in a weekend." "This paper will revolutionize everything." "Just dropped." Most of them are a screenshot and a promise, and I've learned to scroll past.

Every technology I've worked with arrived with someone promising it would change everything. The difference now is the volume and the speed. What I look for is someone who tried the thing on real work and tells you what broke.

"just dropped!!""will revolutionize everything""this guy did it in a day"your filtermost of itgoes nowherescrollpasttried it on real work, said what broke
Most of the feed is shouting. Keep what was tried on real work.

These are the people whose writing I look forward to, and I'm grateful they take the time to share it:

What's different now

Not everything has an old name. In most of the pairs above, the decision stays with me or my code: the hook blocks the push, the CI check fails the build, the schema decides what can be asked. When the model decides which tool to call next, or what to try after a step fails, I don't have an old name for that. That part is new, and I'm learning it the same way I learned everything else.

The other change is what I carry between projects. For years, a new Java service started at Spring Initializr: pick the dependencies, download a zip, build from there. Now that the model writes the code, I rarely start from a zip. I start from a short list I set up before the first prompt: guardrails, one small job per agent, a schema designed from the questions, a goal before any long run, a command for anything I do twice, skills pulled from a shared repository, and a hook for anything that must never happen. The model writes the code fresh each time. The list is what I take from one project to the next.

starter.zipthen
Recipeguardrails firstone job per agentschema from questionsgoal before long runsa command for repeatsa hook for never-againsnow
The zip used to travel between projects. Now the list does.

A list doesn't choose what to build, though. Knowing what's worth building was the hard part in 2004, and it still is. I wrote about that separately.

If you're new to this and it feels as if everyone else got a manual, they didn't. They're trying things and reading each other, the same as you. Some of them write it down, and I'm thankful they do. If there's someone you look forward to reading, I'd love to hear who.


  1. Zheng et al., "When 'A Helpful Assistant' Is Not Really Helpful", Findings of EMNLP 2024 (162 personas, no gain on factual questions); Basil et al., "Playing Pretend: Expert Personas Don't Improve Factual Accuracy", 2025. ↩

  2. Gamma, Helm, Johnson and Vlissides, Design Patterns (Addison-Wesley, 1994), Introduction: "None of the design patterns in this book describes new or unproven designs. We have included only designs that have been applied more than once in different systems." ↩

  3. Tobi Lütke proposed the term in June 2025 and Andrej Karpathy endorsed it; Simon Willison collected both. ↩

  4. Barry Zhang, Keith Lazuka and Mahesh Murag, "Equipping agents for the real world with Agent Skills", Anthropic, October 2025. ↩

  5. Claude Code's /goal keeps working until a completion condition is met; under the hood it is a Stop hook. Boris Cherny, "five tips for running Opus autonomously for hours", June 2026. ↩

  6. Andrej Karpathy, autoresearch, March 2026: the agent edits a training script, trains for five minutes, keeps the change if the result improved, and repeats. ↩