Shipping a Solo-Built Product Using Cloud Coding Agents
Coding agents let solo developers delegate entire tasks, not just suggest code.

Shipping a product alone used to mean trading speed for quality, because one person could not hold architecture, implementation, testing, and iteration at the pace a team could. Coding agents have moved from suggesting code to completing tasks, which changes what a solo builder can hand off, and that is why the old tradeoff is breaking down now.
Shipping alone, then and now
The constraint that defined solo development was never ambition or skill. It was bandwidth. A single developer simply could not simultaneously design a system, write every line of it, maintain test coverage, and respond to feedback at the speed a funded team could manage with five or six people splitting that load. That ceiling held for a long time, and it shaped what solo builders realistically attempted.
What broke the ceiling was not a smarter autocomplete. Autocomplete still required a person at the keyboard, approving each suggestion in real time, holding the plan in their head while the tool filled in syntax. Agents work differently: a builder describes a feature, walks away from the machine, and comes back to a pull request rather than a string of suggestions to accept or reject. That is a difference in kind, not degree.
The clearest illustration comes from a solo-built tarot and Jungian psychology application spanning a large number of files, with three separate LLM models in production, a PHP backend, a web frontend, and an Android app. The developer behind that project reported that writing code was roughly 40 percent of the actual work. The rest was the kind of scope a junior team member would normally own: GDPR compliance documentation, bot-traffic hardening, and the orchestration of development, staging, and production environments on a raw VPS. That is not a tool completing a function call. That is a tool absorbing an entire category of operational labor that previously required either a second hire or a long, tedious weekend.
What the review-and-steer workflow looks like in practice
The shift in capability produces a shift in role. The solo builder who gets the most out of these tools is acting as a project manager and reviewer, not primarily as a typist. The skill that matters most is orchestration: knowing which layer of the system to point an agent at, when to let it run unsupervised, and when to sit down and audit every line it produced.
The workflow itself has a repeatable shape. A builder writes a task description with enough context for the agent to form a plan, lets the agent execute inside an isolated environment, and returns later to a pull request or a diff. From there, the job is to review, correct course where needed, and send the next task. That loop, not the act of writing code, is the unit of solo development now.
This is a different model from pair programming, where a person sits beside the output the whole time, and from autocomplete, where a developer never fully lets go of the keyboard. Delegating to an agent that returns a finished pull request resembles handing work to a contractor who operates while the builder sleeps. The appeal comes with a cost: time that used to go into writing code now goes into reading it. The bottleneck has moved, not disappeared, and any honest account of this workflow has to say so.
That bottleneck is more tolerable when the agent's output is already trustworthy before a human looks at it. An agent that can install packages, start services, and check its own work inside a sandboxed environment returns something closer to a verified draft than a guess.
Giving an agent enough context to work unsupervised
Unsupervised work depends on context, and for a solo builder, the single highest-leverage investment is a project instruction file that stays lean. An agent asked to work without a person present needs enough information to make the same architectural decisions a team member who has been on the project for months would make. It does not need, and should not receive, everything.
The file that works well contains build commands, the project's architectural conventions, naming rules, a list of directories the agent should leave alone, and any domain logic that would otherwise have to be rediscovered from scratch every session. What does not belong is documentation the agent can already read directly from the codebase, or instructions so long that they eat a meaningful share of the context window before the agent has done any work.
Session length matters for a related reason. Agents like Claude Code use auto-compaction on long sessions, summarizing the trajectory of the work so far once the context window fills. That summarization compresses detail, so a task designed to run far past a reasonable window risks losing the thread partway through. The practical implication is to design tasks that complete inside a bounded session rather than assuming an agent will manage unbounded scope indefinitely. Context discipline, not context volume, is what keeps unsupervised runs coherent.
Choosing between Claude Code, Codex, and Opencode for solo work
The useful question for a solo builder is which workload each agent handles most reliably, because the builders shipping the most work run more than one agent.
Claude Code's practical strength appears on tasks that need planning before execution: deep reasoning across a codebase, and architectural refactors that touch many files at once. Developers describe it as the escalation path they reach for when a problem has already defeated a faster, cheaper tool.
Codex is built for a different shape of work. Its strengths are faster execution on tasks that are already well scoped, a lower cost per completed task, and native integration with Linear, GitHub, GitLab (in beta), and Slack. That makes it a natural fit for the background-delegation pattern, where a solo builder queues work from wherever tasks are already tracked. In a head-to-head benchmark across 30 tasks, Claude Code and Codex completed the same number of tasks successfully. Claude Code finished faster; Codex cost less per completed task. The choice between them, on that evidence, is a tradeoff between speed and cost at equivalent reliability, not a question of one tool being better than the other.
Opencode takes a third approach. It uses an AGENTS.md file to hold persistent project context, the same kind of lean instruction file described above, and it integrates the Language Server Protocol to feed the agent real-time diagnostics about code structure. That combination makes it a reasonable choice for a builder who wants open-source control over the agent layer and the flexibility to bring their own model key.
Surveys of working engineers find that a majority still rely on a single agent for daily work, though multi-agent use is growing, typically splitting one tool for deep reasoning from another for rapid execution. The reason to resist settling on just one tool is practical rather than ideological: forcing every kind of task through a single agent sacrifices performance at the seams where that agent is weaker. Locking into one harness also carries a quieter risk. When a better model ships, a setup that is not tied to a single vendor lets a solo builder adopt it without rebuilding the entire workflow around it.
How the agent's environment shapes the quality of what it returns
The infrastructure an agent runs in is part of the quality system that shapes whether the pull request landing in front of a solo builder can be trusted without extensive manual verification. An agent that can install its own dependencies, run the test suite, start services, and check its own output inside an isolated environment produces work that has already been through a review loop before a human ever opens the diff.
That isolation is what makes the difference. An agent running in a clean, sandboxed environment can execute the full test suite, observe what fails, and correct course before returning anything. An agent without that isolation cannot do the equivalent work. Agents on a shared local machine interfere with one another's file state. Agents without network access cannot pull the dependencies a task requires. Agents without a clean operating system layer cannot reliably reproduce the conditions the code will actually run under in production.
The sandboxing techniques behind this vary by platform, from macOS Seatbelt and Linux Landlock, seccomp, or bubblewrap built into CLI-based agents, to full microVMs for cloud-hosted sessions. The common property across all of them is that each agent run gets its own clean compute environment, so whatever one task does to its environment has no effect on the next. For a solo builder running several agents on different features at once, that isolation is what makes it safe to walk away from all of them at the same time.
Triggering agent tasks from where the work already lives
Switching to a dedicated agent interface every time a task needs delegating is its own kind of tax, and agents that can be triggered from Slack, Linear, or GitHub remove it. A solo builder already tracks work somewhere, whether that is a GitHub issue, a Linear ticket, or a note left in a Slack channel as a reminder. Delegating a task from that same surface, instead of opening a separate tool and re-describing the work, removes a context switch that adds up across a day of queuing tasks.
The output side matters just as much as the trigger side. An agent that returns a pull request already linked back to the originating ticket closes the loop automatically, leaving a checked-off, traceable result ready for review with nothing to reassemble by hand.
Multi-agent orchestration tools that coordinate a lead agent with worker agents, routing tasks across Claude Code and Codex with Slack and GitHub integration built in, point to something broader: the layer responsible for triggering and routing work is becoming its own distinct concern, separate from whichever agent actually executes the task. That separation lets a solo builder change which agent handles the work while keeping how the work gets assigned the same.
Breaking large features into parallel agent tasks rather than one long sequential run
The most common failure mode for a solo builder using agents is handing one agent too much scope and letting it run long. Long sequential runs degrade in predictable ways: the context window fills, the agent's attention spreads across too many concerns at once, compaction summaries lose the detail that mattered, and a failure that occurs late in the run can invalidate hours of earlier work that depended on it.
Breaking a large feature into smaller, parallel sub-tasks, each handled by its own agent run with a fresh context window, avoids most of that. Each sub-agent works a well-scoped problem with a focused context. A lead agent coordinating the pieces can keep its own context small because it only needs to track high-level status, not implementation detail. When something fails, the failure stays contained to that one sub-task.
The tarot application is the clearest case for this pattern. Its developer held coherence across a sprawling codebase not by feeding an agent the entire codebase at once, but by knowing exactly which layer to point a given task at and when to do it. That judgment, deciding how to break the work apart, is the actual skill in solo development now. Writing the code is work the agent does; designing the decomposition is work the builder does. The same isolated environment that makes a single agent run trustworthy is what makes running several in parallel safe, since each task executes in its own environment with no state shared between them.
Measuring what the agents shipped so you know what to delegate next
None of this compounds without measurement. A solo builder who cannot attribute output to specific agent runs, specific models, and specific tasks has no way to tell which parts of the workflow are working and which are quietly wasting time.
Without that attribution, basic questions go unanswered: which agent is completing tasks reliably, which tasks are costing more than their scope justifies, and where the agent's output is requiring the most rework before it is safe to merge. Those are the questions that determine what gets delegated next, and they cannot be answered from memory or impression.
The review-time cost named earlier, the fact that agents returning pull requests shifts the bottleneck from writing to reading and merging, turns into a useful signal once it is actually tracked. If review is consistently taking longer than it should relative to the size of the task, the variable to adjust is task scoping or instruction quality, the two things earlier sections identified as the real levers in this workflow. Measurement is what separates a builder who uses agents occasionally from one who has built a system that keeps getting better the more it runs.