Est.

Cloud vs Local Environments for Coding Agents

Persistence, parallelism, and isolation distinguish cloud execution from local coding agents.

Senior Editor, Infrastructure & Runtimes · · 10 min read
Cover illustration for “Cloud vs Local Environments for Coding Agents”
Cloud Agent Environments · October 8, 2026 · 10 min read · 2,269 words

Start a coding agent on a task, then close the laptop and walk away. Whether the work is still running when you open the lid again is a direct test of what kind of system you built, and it exposes the real distinction between cloud and local execution environments for coding agents. A local coding agent runs directly on the developer's machine, drawing on the files, CPU, memory, network, credentials, and development tools already sitting there. The gap between those two setups decides whether an agent can persist after a session ends, run alongside other agents without contention, stay isolated from the host machine, and check its own work against a build that will reproduce the same result twice.

Most teams still talk about this as a single choice, cloud or local, but the architecture underneath actually splits into two independent planes. The control plane decides which model runs, enforces policy, and settles cost. A product can sit on the cloud side of one plane and the local side of the other, and the more interesting architectures on the market today are built exactly that way. The execution plane answers the privacy question of whether a repository gets copied onto remote infrastructure. Cost routing and spend policy belong to the control plane, and they work better when you centralize them instead of leaving them to whichever harness a given engineer installed.

What Local Execution Gives Agents

Local execution hands an agent the most accurate possible picture of one specific machine's state, and for certain classes of bugs that accuracy is the whole game. An agent running locally sees the real repository the developer is working in: uncommitted changes, generated files that never made it into version control, local configuration, a branch that is broken and has not been pushed anywhere. Local execution favors steering; cloud execution favors delegation. If data residency rules prohibit sending code to remote infrastructure, or the environment is air-gapped, or a workflow is tied to specific hardware peripherals, local execution is not a stylistic choice but a requirement.

That accuracy comes at a fixed cost. Local coding agents, including harnesses such as OpenCode (an open-source agent harness, not strictly local), Cline, Continue, and Aider, need the host machine to stay powered on. Close the laptop and the work stops with it. Local development ties a session to a single machine: if that machine is off, the work is unreachable, with no handoff to another engineer, no link to share, and no way to review it asynchronously. That constraint is what the next generation of cloud-based agent environments was built to remove.

What cloud execution unlocks that local cannot replicate

Cloud execution environments unlock four capabilities that a local machine cannot structurally provide: persistence, parallelism, isolation, and self-verification. Persistence means a cloud workspace keeps working after you disconnect and keeps its state between sessions, so an agent can pause and resume instead of starting over in a freshly wiped environment each time. Parallelism means multiple agents can run in independently isolated workspaces at once, without fighting over the same CPU, memory, or filesystem, and cloud tasks can run in the background, including several at a time, while keeping two agents from editing the same files. Sessions stay live and shareable by link rather than bound to one laptop, and the output of a cloud agent task tends to arrive as a diff or a pull request, already in the format a team's review process expects.

Parallelism and self-verification are where the architectural edge of cloud execution becomes hardest to argue with. Multiple agents can run in sandboxed environments at the same time without the resource contention that defines local execution, and each one can check its work against a reproducible build loaded with the team's actual toolchain instead of whatever happens to be installed on one developer's laptop that week. Replicas, for instance, runs each agent task inside its own isolated Linux VM, pre-loaded with the team's dependencies and tooling, which makes the environment functionally equivalent to, and in practical terms often better than, a local developer machine for both verification and parallel execution.

The walk-away test from the opening section is the cleanest way to separate agents that genuinely deliver on this from agents that only appear to. Cloud coding agents keep working after you disconnect, and that single property is what makes them fit for long-running, multi-step tasks handed off outside normal working hours. Locally executing harnesses fail this test, and the honest way to describe that failure is as a capability gap tied to where the execution happens, not a difference in philosophy or taste. A background task that cannot even boot the application because dependencies are missing or the environment is misconfigured is an error log waiting to be read, not a worker quietly making progress. The environment an agent runs in carries as much weight as the model driving it. Configuration deserves the same scrutiny as model selection.

How isolation and environment configuration shape agent reliability

Isolation inside a cloud environment does not happen automatically. It comes from specific configuration decisions, and those decisions decide whether you can trust an agent's output or only hope for the best. A working cloud coding agent depends on five components functioning together: the language model, running on the provider's own infrastructure and separate from the execution environment itself; the tool layer, covering shell commands, file reads and writes, Git operations, package installs, and test runs; the execution environment, the OS, development tools, repository, filesystem, and allocated compute; secrets and credentials for model access, Git repositories, APIs, and package registries, available to the workspace without ever being exposed in the codebase; and state and persistence, the files, dependencies, and Git history retained across sessions.

Standard containers share the host kernel, so they are not enough for AI-generated code that may try privileged operations it has no business attempting. You can address this with microVMs, with gVisor, which implements a user-space kernel, or with hardened containers, and each trades off overhead against security guarantees differently. Network configuration is the control that matters most in practice: default-deny egress, where every outbound connection is blocked except destinations explicitly allowed, is the single most important network setting for an agent sandbox, since it limits what a misbehaving or compromised agent process can reach. Ephemeral environments give every task a clean slate and prevent one run's leftover state from contaminating the next. Persistent environments let an agent accumulate context, installed dependencies, and Git history across sessions, trading a clean start for continuity.

The environment boundary itself is what guarantees reliability in that setting, not a review step added afterward. That is precisely why teams building agent infrastructure at scale reach for isolated virtual machines rather than shared containers when the work has financial or compliance stakes attached to it.

How the leading harnesses split across the local-cloud execution boundary

Every harness a team adopts sits somewhere on both the control plane and the execution plane, and that position determines which classes of tasks the team can actually hand off to it. Separating the two planes clarifies how a team stays flexible about which agent it uses while still centralizing policy in one place: the control plane governs which agent gets triggered, what spend limits apply, and what approval workflows gate a merge, while the execution plane, cloud or local, is where the agent's work actually happens. Replicas implements that split by running agents inside isolated cloud VMs, triggered from the workflows a team already uses, Slack, Linear, GitHub, or GitLab, so the team is never locked into a single agent regardless of which harness fits a given task.

Replicas runs coding agents, whether that is Claude Code, Codex, Opencode, or another agent the team selects, inside isolated, fully configured Linux VMs on remote infrastructure, with each run sandboxed in its own VM and pre-loaded with the team's dependencies and tooling. It passes the walk-away test by construction: the work continues after the developer disconnects and comes back as a pull request, a recording, or a reply, without requiring anyone to babysit a terminal window. Replicas also provides attribution down to the source, harness, model, and credential behind every minute an agent runs, giving teams the control-plane function to measure and adjust cloud agent spend. Because it stays harness-agnostic, a team can swap in a better model as one appears without re-engineering the environment or the integrations built around it.

Claude Code works as a terminal-native agent, and it offers several distinct execution surfaces. You can close the laptop and its cloud sessions and scheduled Routines keep the work going, while background agents return output asynchronously once a run finishes. The control plane here is the vendor's cloud, but the execution plane depends on which surface a developer picks, a local terminal, a desktop worktree, or a hosted session, each with different consequences for whether the walk-away test passes.

Codex leans cloud-first, with a local CLI available as an alternative. Cloud tasks run in OpenAI-managed isolated environments, each starting from a pre-published, reusable environment where the repository and dependencies are already prepared, and the agent edits code, runs checks, and returns a diff, a strong fit for independent, reviewable work with clear acceptance criteria. The local CLI and IDE extension, by contrast, operate against a local workspace, and local schedules require the computer to stay on, the project to remain on disk, and the desktop app to keep running. You can delegate natively to Codex from Linear, by assigning an issue or mentioning it in a comment, and that integration is reported to produce solid, consistent quality.

Opencode splits the difference differently: a local server that remote clients can drive, with remote client access also supported. Because execution stays local, the machine has to remain on, and Opencode does not pass the walk-away test for long, unattended tasks. That makes Opencode a reasonable fit for teams that want remote access to an agent process while keeping execution on infrastructure they control.

Set against the walk-away test, the split is stark. Replicas passes by design, since its cloud VMs keep running regardless of the developer's connection. Claude Code passes on its cloud sessions and scheduled Routines, but not on its local terminal surface. Codex passes on cloud tasks, but its local CLI and local schedules depend on the machine staying on. Opencode, running locally even with remote client access, fails the test.

Matching environment type to the tasks teams need to delegate

The environment decision should follow the shape of the task's dependencies, not whichever tool happens to already be installed on a given machine. Cloud environments fit independent issue fixes that come with clear acceptance criteria and tests, several unrelated backlog items that can run in parallel without touching the same files, cleanup after review comments where the output naturally becomes a branch and a pull request, and refactors, documentation updates, security review follow-ups, and migration work after a schema change, anything that can be described as a repo-scoped change with a clear definition of done. Local environments fit bugs that only reproduce against a local database, a private VPN, a browser profile, or a mock service, work built around uncommitted experimental files where local state is the only source of truth, high-friction judgment loops where a developer wants to watch and steer execution directly, kernel-level work that needs direct hardware access, and any environment bound by data residency rules that prohibit sending code to remote infrastructure. The working rule: route independent, reviewable work that benefits from running in parallel to cloud tasks, and keep high-context work that depends on the current branch, local credentials, or close supervision on local agents.

One mistake undercuts this framework more than any other: treating "background" as a quality score in itself. A background task only earns that label if what it returns is trustworthy enough to review without extra cleanup. If the environment is misconfigured, dependencies are missing, or two agents end up editing the same files, the background task does not quietly save time, it produces background confusion that someone still has to untangle by hand.

Most teams that get this right are not choosing cloud or local as a permanent stance. They are running a handoff loop. Local work discovers the shape of a problem, then cloud tasks execute bounded branches of that problem once the shape is clear, and local verification decides what actually survives code review. Complex reasoning work, debugging, architecture decisions, multi-file changes, tends to suit cloud agents, while high-frequency routine work, completions, tests, documentation, suits local agents better, and that split reflects where the industry's hybrid architecture is converging. This split in task allocation changes what engineers spend their time doing in orchestrating long-running systems of agents: fewer hours writing every line by hand, more hours orchestrating, with the environment infrastructure underneath that orchestration keeping it reliable instead of chaotic.

Measuring whether the environment choice is working

An environment decision made once and never revisited is not a decision a team can defend. You need visibility into what agents are doing, how long each task takes, and what quality the resulting diffs carry once a human reviews them, to measure whether cloud or local execution is actually paying off. Attribution down to the source, the harness, the model, and the credential behind each run turns agent spend from a line item nobody can explain into something a team can actually audit and adjust. Without that visibility, the choice between cloud and local is no longer an engineering decision; it is a guess that happens to run in production.

More in Cloud Agent Environments