~/writing/how-im-building-my-agentic-workflow

How I'm building my agentic workflow

I think it was one of Theo Browne’s videos that pushed me to take my setup more seriously. I don’t remember the exact wording, but the idea stayed with me: don’t copy someone else’s setup. They have different needs. Experiment with your own.

At the time, my workflow was mostly asking an agent to grill me about a feature, turning the conversation into a spec, and implementing it. That already worked reasonably well. But I started noticing how many things I kept correcting, repeating, or checking manually.

Those became opportunities to change the workflow itself.

Most of this experimentation happened while building Rhea. This post describes where my setup is now, why I made certain choices, and which parts I still don’t trust enough.

Choosing models for different jobs

I use both Pi and Claude Code. If my Claude subscription were supported in Pi, I’d prefer to do everything there. I like Pi because it’s minimal, has a small system prompt, and lets me extend it with plugins and my own tools.

But I still want access to Anthropic’s models.

In my experience, they were better at frontend work, and I started noticing that I generally preferred the structure of the code they produced. Early in a project, I would lean on them to establish patterns, then use lighter GPT models to continue implementation.

GPT-5.6 models were very good at solving difficult problems. Sometimes, though, the implementation felt like brute force. It solved the problem, but I didn’t always like what it left behind. Repeated across enough features, that could turn the codebase into something I wouldn’t want to maintain.

That’s my experience with those models on this project, not a permanent ranking.

Subscription limits also influence the setup. If I use Fable for every discussion, implementation, and review, I burn through its allowance quickly. Having it mostly orchestrate work while other models handle specific tasks helps me spread usage across subscriptions.

I’m trying to get controlled software delivery out of the models available to me, without spending the most limited capacity on every small task.

Explore before asking what to build

One distinction I’ve become more deliberate about is exploration versus grilling.

Exploration is about understanding the existing system. Grilling is about figuring out what I actually want to build.

Suppose I want a user preferences module. That’s an idea, but it doesn’t tell me which preferences belong there, how they relate to existing behavior, or what the implementation should look like.

I want the agent to explore first, because what it finds can completely change the feature we’re about to discuss.

Agents often focus too narrowly on the current ticket. They find somewhere to put the code, satisfy the acceptance criteria, and move on. They can miss existing conventions and patterns, but the bigger problem for me is missing opportunities to improve the system’s structure.

If we keep adding locally reasonable features, we can still end up with a messy system.

I ask for schemas, pseudocode interfaces, and call stacks to understand the proposed design. That gives me something concrete to judge. Should this use an existing service? Is there a useful shared abstraction here? Are we making one component bigger when we should introduce another building block?

I don’t have a mechanical rule for those decisions. Some of it is taste. I need to see how the system works before deciding what should change.

Then we grill the idea.

Use the conversation to build the specification

Grilling helps me discover what I mean by a feature.

For substantial changes, I want the agent to question assumptions, present options, and explain the consequences. I’m not doing this for every spacing adjustment in the UI.

Sometimes I accept a recommendation immediately. Sometimes I need an example or a simpler explanation. I use /bro and /eli5 because I don’t want to agree to a design just because the explanation sounds convincing.

The result should be a specification based on decisions we actually made.

For larger work, that becomes tickets with clear problem statements, goals, dependencies, and enough interface detail to give implementation agents a good starting point.

The aim is to remove avoidable guessing. If something turns out differently from our assumptions, I want the agent to say so.

Turn repeated instructions into skills

I noticed that I was starting implementations with the same few sentences.

Use these models. Follow the Effect patterns. Delegate UI work to Claude Code. Validate before review. Start a fresh reviewer after substantial fixes.

So I turned those instructions into an /implement skill.

There’s no elaborate theory behind that decision. I was repeating myself, and a skill removed the repetition. It also gave me a documented workflow that I could change deliberately.

I prefer short constraints to long scripts written in prose. The agent should still work out how to complete the task.

Skills also help with context management. Git and PR instructions don’t need to occupy the beginning of every discussion about architecture. They can be loaded when the agent is about to do that work.

Keep delegated work visible

I use Herdr to coordinate agents across Pi and Claude Code.

An orchestrator can start implementation agents in separate panes, assign work, collect their reports, and launch reviewers. Task worktrees keep the work separate, and I changed the workflow so the orchestrator moves into the relevant worktree too.

The visibility matters to me. I can inspect an agent, see whether it’s stuck, and intervene instead of relying entirely on the orchestrator’s summary.

I also want tabs named clearly and closed when they’re no longer useful. A terminal full of abandoned agents doesn’t give me more control.

Reviews help, but they can also create work

My workflow asks implementers to validate their own work before handing it to independent reviewers.

I eventually separated review into different jobs. One reviewer looks at structural code quality. Another checks repository standards and whether the implementation matches the specification. After substantial fixes, I want fresh reviewers.

This catches useful problems. It also creates problems of its own.

Some features went through many rounds of review and fixes. Agents kept finding increasingly unusual cases, and delivery kept moving further away. I wanted to test ideas and deliver something useful, not perfect every corner before anyone used it.

So I added instructions to allow non-blocking work to be deferred.

That policy is still too vague. Right now, agents have considerable freedom to decide what deserves fixing and what can wait. I need better guardrails there.

There’s another risk: every fix round can move the implementation further from the original specification. We’re supposed to catch that in subsequent reviews and my own inspection, but it’s another thing to watch.

More review isn’t automatically more confidence.

I still need to understand the system

I’m not reading every line of generated code anymore.

I do want to understand the important flows, the rules the system follows, and whether the architecture and coding standards I care about are being preserved.

Plannotator’s code-review guide helps here. A guide to what changed and why gives me a way into the diff without starting at the first changed line and reading everything in order.

That has helped me notice problems that passed automated checks and agent reviews. In one case, I found that an implementation wasn’t following the Effect patterns I expected, despite the workflow telling agents to use them.

Instructions being present doesn’t mean they’re being followed.

The most confidence still comes from switching to the branch, running the application, and trying it myself.

What still needs work

Too much of my setup currently relies on agents remembering to do the right thing.

They should run formatting, linting, tests, and other checks. But asking them to do that isn’t the same as enforcing it. Hooks and automated gates could make those steps harder to skip.

I’d also like more lint rules for recurring code patterns I don’t want. If a tool can detect a problem reliably, I’d rather encode it than keep explaining it in reviews.

Application verification is another gap. I want agents to run what they built, exercise the changed flows, and attach useful evidence to the PR. Screenshots, and sometimes video, would help me review the result and understand what was actually tested.

Longer term, I think more of the workflow should be code.

Herdr gives me a way to coordinate agents across different applications today. But starting the orchestrator, managing execution, running checks, and enforcing completion conditions could be controlled programmatically instead of described in instructions.

I haven’t built that yet.

For now, I’m continuing to watch where the process fails. Sometimes the fix is a skill. Sometimes it’s a different model, a better spec, or a check that shouldn’t depend on an agent’s judgment.

That’s the part of this setup I would encourage someone else to try: notice what isn’t working in your own process, and change it.