The Code/X ArchiveView on X
Tomas Vykruta

@tvykruta

1/ Yesterday I shared a breakthrough.

I spent years writing C and C++ at Microsoft and Google. I took the 10 design principles I learned back in the '90s and turned them into hard rules for Claude Code.

Today, I'm presenting the results.

My app is 100% AI-written with Opus. Before the rules, one feature ate 5 days, 49 PRs and ~$640 in tokens. It was still full of bugs.

After the rules, Claude rebuilt it into some of the best code I've ever read. $7 → $1 per PR. 9× fewer tokens. Every PR since merged on the first try.

Bookmark this. My latest AGENTS.md rules are at the end.
27451.1K2.9K
Tomas Vykruta

@tvykruta

2/ The staff-engineer architecture review was a tragedy:

• Pricing logic spread over 32 files
• 10 formulas copied 40 times
• 12 copies disagreeing: same item, two prices on two screens
• One 2,800-line file running every page
• 16 of 52 rule checks passing

Textbook answer: rewrite from scratch.
102312
Tomas Vykruta

@tvykruta

3/ And the god code, actual nightmare for a code reviewer:

• Main app file: 2,916 lines, 62 functions • One function 321 lines long
• The same 200-line feature written twice, once per platform
• A pricing file that grew from 757 to 1,950 lines in 3 days, parsing, pricing and rendering all at once
10157
Tomas Vykruta

@tvykruta

4/ Now the cost. Over 5 days, 49 of my 166 PRs (30% of all engineering on the app!) were spent trying to fix one pricing problem.

The cost:
• 50+ agent runs, 23 agent-hours
• 11 runs on a single icon
• +6,936 lines written, −1,979 ripped out
• One 3-hour rewrite across 89 files, thrown away entirely

All those tokens produced garbage. The problem was never fixed. Each patch made it worse: more copies, more places where they disagreed, more ways to break.

In staff-engineer terms, that's negative velocity: every PR made the next one more expensive.
10135
Tomas Vykruta

@tvykruta

5/ The fix wasn't a better prompt. It was a new process.

I encoded my 10 principles as hard rules in AGENTS.md. The most important one changes how Claude works on every task: *two steps, never one.*

Claude is NEVER allowed to implement a change until it has refactored the code first.

Step 1: Refactor the existing code to make room for the change. Behavior stays identical, and the green unit tests prove it.
Step 2: Only then implement the change, on the clean structure.

No refactor, no change. Both in one diff is a rule violation.

An agent that goes straight to step 2 patches around the mess. That's where the 49 PRs came from.

For the complex pricing system, step 1 started with an staff-engineer audit before any code was touched ($11, and it found 7 bugs nobody had reported). Then 11 behavior-preserving refactor steps consolidated everything into one module. Only then did new work land on top.
103140
Tomas Vykruta

@tvykruta

6/ A new pricing module was born. Some of the best code I've ever read.

1 package, 14 small modules, 1 rulebook: constants → market → typical price → trend → estimate → portfolio value

One-way dependencies. No cycles, no inheritance, no web code.

Principle-by-principle review ↓
10167
Tomas Vykruta

@tvykruta

7/ Separation of concerns ✅
• Price code outside the module: 76 → 0
• Math in templates: 16 → 0

Cohesion ✅
• Files holding price logic: 31 → 1 package
• Imports buried in functions to dodge circular dependencies: 60 → 1

DRY ✅
• Formula copies: 40 → 10, one each
• Hand-rolled formatters: 11 → 1
• Constants outside one file: 17 → 0

Encapsulation ✅
• Back-door modules: 5 → 0
• Outside code touching private internals: 0

KISS ✅ conflicting versions of the same rule: 3 → 0
Composition ✅ 0 inheritance; 22 of 26 classes are plain data

Contracts ✅ 0 web imports; all of it tests without the app

YAGNI ✅ −4,803 lines. Unjustified 5× multipliers and invented prices → deleted

Open/closed ✅ change how prices are sampled in 1 place, not 7.

Checks guarding all of this: 3 → 7. Break a rule and a check fails before it merges.
102120
Tomas Vykruta

@tvykruta

11/ Claude learned to code from the internet. Garbage in, garbage out.

So I gave it the principles I used writing C++ at Google under Jeff Korn.

Same Claude. Same codebase. New rules. Astonishing results.

7× cheaper per PR.
5× faster PR completion
9× fewer tokens.

* These are measured, not estimated: a before/after analysis of the codebase and the agents' own session logs.

And I did it on Opus. One $11 planning pass used Fable; every refactor step and every feature since ran on Opus, the cheaper model. On a clean codebase I don't need the bigger model. Opus ships new features in one shot, and every PR since has merged on the first try.

The rules are below. Steal them. ↓
10267
Tomas Vykruta

@tvykruta

12/ How to apply this to your own codebase. Keep it simple:

Put the rules below in your AGENTS.md (or CLAUDE.md).

Pick the system that hurts most. Ask Claude for a staff engineer architecture review of just that system against the rules. Numbers, not opinions: every file holding its logic, every copy of every formula, every place the copies disagree.

Ask for a refactor plan: small steps, each one behavior-preserving, one PR each. No new features in the plan.

Run the plan step by step. Tests green after every step. If a step changes behavior, it's a bug in the refactor.

Add a check that fails if the mess comes back (a copy, math in a template, logic outside the owner).

Ask Claude for a second staff-level architecture review of the same system. Compare it to the first, pattern by pattern, with the same numbers. Anything not fully compliant goes back to step 3.

Only then build the feature you wanted in the first place.

Then pick the next system.
102155
Tomas Vykruta

@tvykruta

13/ My AGENTS.md, straight from the repo. The design principles (every file):

::: [

Constraints, not a checklist; when two conflict, pick the lowest future cost for this repo and say so in the commit. AGENTS.md invariants 10–12 are the hard form of these.

- Separation of concerns: domain (services/<domain>/) · persistence (models.py, alembic/) · presentation (templates, static/) · routing (pocket/, app.py). Name the one concern of the file you edit.

-Encapsulation: public contracts only; read another module's tables and caches through its functions.

-Cohesion / coupling: one rule change touches one module.

-DRY: grep before writing logic; one home per rule, threshold, format or schema fact. Do not abstract coincidental similarity.

-KISS / YAGNI: simplest working shape; function over class; no speculative hooks, flags or frameworks.

-Single responsibility: if you describe it with "and", split it. Names say intent; comments say why.

-Depend on contracts: domain code takes and returns plain values; never imports Flask, request or templates.

-Composition over inheritance · open/closed only where change has happened twice · Demeter (no a.b.c.d) · fail fast (validate at edges, never swallow errors) · optimize for deletion · boring tech.

The hard invariants (never violate):

10\ One owning module per domain — each business domain's logic (pricing, valuation, pedigree reading, sharing, imports, analytics…) lives in one module with its rules doc; everyone else calls it. Pricing: services/pricing/ + docs/PRICING_RULES.md. Consolidate a scattered domain before adding to it. Extend the domain's existing module; create a new one only when you can say why the old one cannot own it. A new rule goes in its domain's owner, not in the first feature that needs it; a function-level import to dodge a cycle means the logic is in the wrong module.

11\ Never duplicate logic — second use of existing logic: (1) move it to a shared module, (2) switch the original caller, tests green, no behaviour change, (3) then build the new use. First grep for the expression (the arithmetic, the format string, the threshold) and list every copy; the move switches them all or the commit names each one left and why. A new helper beside old copies is one more duplicate. The move keeps each caller's exact results (guards, rounding, clamps); any behaviour change is its own commit. Same for schemas (one fact, one column) and UX (one partial per repeated piece).

12\ No business logic in rendering — templates, JS, routes and view builders only display values; they never compute prices, rules or classifications. Arithmetic or rules on business data there is a bug; move it to the owning module. Who sees a value is decided in Python, not by a template if.

How to work:

Refactor first, then change: a behaviour-preserving refactor with tests green (plus an output dump for pricing-sized domains), then the change. Never both in one unverifiable diff.

Gates, not promises. A prompted "never" alone does not protect the SoT; a rule that matters is enforced by CI or a hook.
] :::
49225972
End of thread