Skip to content
Notes

Longread · AI agents

How I run nine projects through Claude Code without being a programmer

On July 7 I asked an AI why I need Homebrew and what wrangler is. Three months later it works with me on a client's growth, a site factory at my job, a network of content sites and a few products of my own. This is the system I built around it — and the failure behind every rule.

Max Kiriienko
Tech Lead SEO & Marketing · 13 min read

Читати українською

I’m an SEO and marketing lead. Ten years of strategy, analytics, content and teams — but not code. I had built a couple of small sites with another AI editor before, but I had never set up a developer’s machine.

On July 7, 2026 I had just bought a Mac. I asked Claude what to install to build websites with GitHub and Cloudflare. Then I asked two questions any developer would smile at. Does Homebrew go first, with everything else on top of it? And what is wrangler — isn’t that something inside Cloudflare, so why would I need it?

The note the agent saved about me that day is blunt: a beginner who didn’t know what wrangler is or why Homebrew is needed. Homebrew was the one thing it couldn’t install for me, because it asks for the Mac password. Everything else it installed by itself in under four minutes.

Three months later the same agent works with me across nine projects. Growth, analytics and SEO for happymonday.ua, a Ukrainian career platform. A site factory at my job that builds, deploys and monitors a network of sites on Cloudflare. A small network of content sites of my own. A houseplant encyclopedia in English and Ukrainian. A small tools site, an AI content app and this site. And two projects with a friend: a Notion graph app and a dictation app for Mac.

This is not a “50 prompts” post. Prompts didn’t make it work. What made it work is a small system around the agent: one rules file it reads every time, playbooks instead of instructions, a memory where every rule is written down with the failure that caused it, and a style guide that makes it write for me instead of for engineers.

The numbers after three months

~1,750

messages I sent, most of them dictated

~300K

words in those messages

~150

memory notes, over 50 of them lessons

7

skills the agent follows

I counted. Between July 7 and October 3 I sent about 1,750 messages. The median message is about 620 characters long; around 280 of them are longer than 2,000. A third contain “uh” and “um”: I talk, I don’t type. About 950 messages came with screenshots — some 2,500 images in total.

Two numbers tell more than the rest. I switched the model by hand about a hundred times. And my work project has lived in one main conversation since July 10: it outgrew the agent’s working memory and was compressed 43 times. One project, one main conversation, three months. That only works if the important things live outside the conversation — in files the agent reads again every time.

Layer one: a global CLAUDE.md the agent reads every time

On August 1 I complained, again, that every new session I had to re-explain keys, SSH, GitHub and Cloudflare. That evening the first version of a global rules file appeared: CLAUDE.md in my home folder. Claude Code reads it at the start of every session in every project. Today it is just over 200 lines long. Right under the title it says: keep this file compact.

It holds five things:

  • Identities. I have four: my job, a client, my affiliate sites and my personal projects. Each one is a set — one Google account, one Cloudflare account, one GitHub. A project always works under one identity, and the agent never mixes them.
  • Where the keys live and how to read them without printing.
  • How I want work done — nine rules, most of them written after a failure.
  • The project standard: every project has a passport and one backlog.
  • What never happens without my “yes”: deleting anything, and buying a domain or anything else that costs money.

Identities sound like bureaucracy until the first accident. On my machine the default GitHub address points to my work key. Without the map, nothing stops the agent from pushing personal code with work credentials. Now personal repositories go through a separate host, and the rules file says so.

Keys deserve their own paragraph. I don’t paste them into chat anymore. They live in KeePassXC, the master password sits in the macOS Keychain, and the agent reads a key with one command and never prints the value. When I get a new key, I drop it into a plain text file in the project root. The agent moves it into the vault, checks it with a live request, and only then empties the file. Empties — not deletes. That word cost me a token, and I’ll come back to it.

Layer two: skills — playbooks instead of prompts

On August 6 I noticed something unpleasant. In one project I had set up backups. In another I hadn’t — and I was sure I had. There was no shared base of practices: each project knew only what I had said in its own conversation.

The same day we introduced playbooks (Claude Code calls them skills), a “method of work” block in the rules file, and a passport plus a backlog in every project. The first scan against the new standard immediately found holes: two projects had no database backup at all, my personal site had neither analytics nor a repository, and some keys still lived in text files. It also found that I had asked the same question in two projects four days apart: how do I set up a Google service account?

The scan got one thing backwards, too. An old line in my own rules file said to skip Tag Manager and use the Google tag directly, and the agent wanted to make it a hard rule. I corrected it within minutes: we always install Tag Manager. The agent then checked the code: seven sites across my projects already ran on it. A playbook has to describe what I actually do, not what one old line says.

There are seven skills now:

  1. Launch a new site — domain, repository, deploy, analytics, Search Console, indexing, backup.
  2. Set up analytics — Tag Manager on every site, one service account per identity, checked on the live page.
  3. Audit backups — a backup counts only after a tested restore.
  4. Audit access — find every secret, move it to the vault, rotate what leaked.
  5. Audit a project against the standard and fill in its passport.
  6. Logic audit — the bugs that tests don’t catch.
  7. Which Ahrefs metrics to trust — and which to ignore.

A prompt is a wish. A playbook is a checklist with checks the agent can run itself. “Set up analytics” means nothing. “One container per site, tags only inside the container, confirmed by the network requests on the live page” is something an agent can do and prove.

The other half is the passport. Each project has a short file with its identity, how it deploys, where its analytics and backups are, and a link to its single backlog. Anything we discussed and didn’t do goes into that backlog. Not into a chat, not into a second document — one file per project.

Layer three: memory, where every rule has a scar

The agent keeps a memory between sessions: about 150 notes across my projects. More than 50 of them are lessons — something went wrong, here’s the rule. I like the format so much that I now distrust any rule that comes without a story. Here are some of them.

WhenWhat brokeThe rule since then
Aug 7Deploys seemed to take up to 20 minutes. A deploy took about ten seconds; a slow generation step was glued into the same command.Estimate how long a command should take before running it.
Aug 7The agent saved a token truncated, read undefined as success and deleted the original file. The token was gone.Check with a live request first. Then empty the file — never delete it.
Aug 14Background tasks sat on "No output yet" for nine minutes. The work itself had taken about a minute.Progress must be visible to the human, line by line.
Aug 15A git history rewrite wiped 31 brand-book frames and 156 cover designs off the disk — right after "nothing is deleted from disk".Nothing is deleted without an explicit "yes" for each deletion.
Aug 25Three new checks passed without checking anything.A check counts only if it fails on a deliberately restored bug.
Sep 21The agent judged a page from a thumbnail: "structurally fine". At full size, 12 of 56 cells were empty.Judge layout by measurements or at full size.
Sep 22"Done" — while the code wasn't in production.Done means deployed and checked with a request.

Read the right column again and you’ll see one sentence repeated in different words: the agent checks the code, not the result. That sentence now sits in the rules file as a method: check the result, not the code. An audit that lies is worse than no audit. Fallbacks must be loud.

Some scars repeat. On September 26 another secret was lost: the value arrived without a line break, and the helper that saves keys printed “✓” anyway. Now the helper reads the entry back and compares. Same lesson, one level deeper.

Layer four: teaching it to write for me

The agent was capable and hard to read. It used names from the code, invented its own terms, and answered with a work log instead of an answer. The share of my messages that complained about clarity grew, by my rough count, from 6% in July to 15% in August and 28% in September. The more complex the system became, the more expensive it was to understand it.

My messages that complained about clarity

Share of messages complaining about clarity: 6% in July, 15% in August, 28% in September.

Share of all my messages that month, by my rough count. Source: my Claude Code history, Jul–Sep 2026

So the agent went through my history and collected 100 pairs of “its answer → my complaint”. The top problems were jargon and English terms I never used (28 cases), names from code (27), a work log instead of an answer (25), self-invented terms (21), walls of text (16) and questions three words long (12). From these pairs came a style guide: the first sentence answers the question; things are named the way I see them on screen; numbers come with a unit and a meaning; every question to me makes sense without context and includes a recommendation.

What 100 “its answer → my complaint” pairs were about

Jargon 28, names from code 27, work log instead of an answer 25, self-invented terms 21, walls of text 16, three-word questions 12.

One pair can count in more than one group. Source: my Claude Code history, Jul–Sep 2026

Before

The gate compares class names, not geometry. A level-two gate is suggested: a Playwright render…

After

The automatic check only looks at block names, not at how they look. I suggest a second check: open the page in a browser and catch overflowing text and empty sections.

One of the pairs from the style guide. My reply to the "before" version was: "A level-two gate? I don't understand."

Then we tested it blind. Two agents rewrote seven real answers I had complained about: one was told to “just make it clearer”, the other followed the style guide. A third agent judged them without knowing which was which. The first version of the guide won only 4 of 7. After fixes, it won 5 of 7, lost 1 and tied 1. Clarity on first reading: 4.1 versus 3.1 out of 5.

The judge’s comment was the real finding: rephrasing doesn’t help if the writer hasn’t drawn the conclusion and answered the reader’s next question. Clear writing is clear thinking that has been finished.

My own first reaction to the report was useful too: “I can’t tell whether you solved the problem or just said you did.” So we added 35 short before-and-after pairs. Show, don’t claim — that applies to agents as much as to people.

Bad writing didn’t stay in my chat. On September 21 the agent wrote a 2,600-word status document and four task comments in one day on a client project. I told it this read like complex technical documentation, and the client’s team said the same that day. On October 1, in my work project, it passed me five questions from another agent almost word for word. None of them said which button it meant or what was happening now. I couldn’t make sense of any of them. To four of the five I answered: just make it logical.

Since October 3 the style guide covers every text a person will read: replies to me, tasks and comments in the team’s tracker, Slack messages, documents, statuses, even buttons and notifications in our tools. Site content is not covered: it keeps its own voice. Questions have their own rule. The agent decides whatever has a logical answer and tells me in one line. It asks me only about real choices: where it is on the screen, what happens now, what changes after “yes” or “no”, and what it recommends.

How I talk to it

I dictate. My messages are long and mixed: a question, a decision, a new task and a complaint in one breath — often while the agent is still working on the previous thing. About 180 of my messages were sent while it was busy.

That has a cost. On August 12 I said: “I dictate something, you answer part of it, and the rest gets lost.” We checked: from 32 messages over two days, about 15 items had vanished. The rule since then is simple. Every message is split into items, and each item is either done, written into the backlog, or explicitly rejected.

Six weeks later we audited nine busy days of my work project: 660 items from my messages. 403 were done, 145 were in the backlog, 21 were cancelled — and 91 had gone nowhere. That last number is why the rule exists, and why it has to be checked, not assumed.

Where 660 items from my messages ended up

660 items: 403 done, 145 in the backlog, 21 cancelled, 91 went nowhere.

Source: nine busy days of my work project, Sep 15–23, 2026

When 50 subagents is the wrong answer

Claude Code can run dozens of subagents in parallel. It feels like power, and the agent itself will happily turn a small check into a hundred agents. Some numbers from my own history:

  • 49 agents launched to answer a question whose answer was already in a file written the day before.
  • An analytics audit found 31 findings, and a script attached three skeptics to each — 93 agents and about 6.5 million tokens planned. I stopped it at 31 verdicts and about 2.2 million tokens. In my words then: “I’m not against big research. It just has to make sense.”
  • One evening, an audit of 24 + 7 agents used up a whole week’s limit of the flagship model.

The rules now: name the number of agents and the token estimate before launching. Around ten agents per task at most. Don’t send agents to verify what one query can answer. Run wide fan-outs on a cheaper model, keep synthesis on the flagship — and never switch to a weaker model silently.

Auto mode: how much autonomy I give it

In July the agent could edit files on its own, but asked me before most commands. Since August I work in auto mode: the agent acts on its own, and a built-in classifier blocks risky actions. It has stopped re-uploading icons to more than a hundred live sites, merging branches written by other agents, and a mass deletion in a live account — even after my explicit “yes”.

We didn’t look for a way around it. For anything shared, public or irreversible, the agent now prepares a script with a dry run and gives me one command to run myself. Purchases and anything that costs money happen only after my “yes” with the price in front of me. Autonomy for edits and research; human hands on everything that can’t be undone.

Two questions people ask me

Can you use Claude Code if you can’t code?

Yes, if you can say what you want and check the result. I still can’t write the code myself. What I do is set rules, ask for evidence and look at the result: the live page, the real numbers. The skill that matters here is managing, not programming.

Where does the global CLAUDE.md live?

In your home folder: ~/.claude/CLAUDE.md. Claude Code reads it in every project. Each project can have its own CLAUDE.md in its folder on top of that. Mine hold a short passport: which accounts the project uses, how it deploys, where its backlog is.

If you’re an SEO lead starting tomorrow

  1. Write one rules file before the first project: accounts, where secrets live, what “done” means, what never happens without asking.
  2. Never paste keys into chat. Use a password manager the agent can read from.
  3. Turn every repeated task into a skill with checks, not a prompt.
  4. Keep one backlog per project. Every item you dictate ends up done, in the backlog, or rejected.
  5. Write every failure down as a rule — together with its story.
  6. Ask for evidence: the live page, the real data, a request. Not “I checked the code”.
  7. Count agents and tokens before big runs.
  8. Teach it how to write for you, and test that blind.

None of this required me to learn programming. It required me to manage the agent the way I’d manage a strong, fast and overconfident new hire: clear rules, written standards, and no trust without evidence.

About the author

Max Kiriienko

Tech Lead SEO & Marketing from Ukraine. I design growth strategies and build the pipelines, tools and teams that execute them.

Read next