Trend shift

Companies Build Their Own Coding-Agent Control Layers

Coinbase, Shopify and Ramp are building internal coding agents alongside tools such as Claude Code. The differentiator may be shifting toward codebase access, credential governance, sandboxes and observability.

Coinbase, Shopify and Ramp are extending coding agents beyond purchased general-purpose assistants into internally built layers connected to codebases, collaboration systems, credentials and execution environments. This does not show companies abandoning Claude Code or similar products, but it suggests differentiation may be moving toward internal workflow integration and control planes.

Internal agents enter daily engineering work

The Information reports that Coinbase rolled out its internal coding agent Forge to all engineers, alongside tools from Anthropic and OpenAI. Shopify and Ramp are also building internal tools. Coinbase had already made Cursor, Copilot, OpenCode and Claude Code available to engineers, indicating that internal development is not simply a response to missing external products but a push to fit agents into company-specific environments.

The moat is connection to internal systems

Shopify Engineering says its Slack-native agent River had co-authored one-eighth of merged pull requests at launch. River runs on Aquifer, Shopify’s internal platform for persistent sessions, sandboxes, credential brokering and observability. Coinbase’s Mux figures likewise show that agent value can be measured through engineering workflow outcomes, though Coinbase cautions that early-adopter selection bias affects productivity comparisons.

Procurement may become a layered model

Once agents are embedded in permissions, code review, deployment and audit workflows, external models and internal control layers can complement rather than replace one another. The strongest countercase comes from a peer-reviewed ACL study: on 105 real multi-file C/C++ tasks, the best agent configuration produced answers that were both functionally correct and secure only 23.8% of the time. More PR output alone therefore does not establish safe replacement of human review.

What to watch next

Watch for whether more companies disclose engineering coverage, pull-request or deployment scope for internal agents; whether these systems expose interchangeable interfaces to external models; and whether long-run comparisons show better security, rework and review outcomes. The claim would weaken if deployments remain limited to early adopters or fail to demonstrate quality gains.

Sources