The complete history of new features, changes, and upgrades shipped to help you become a better Outcome Developer.
Anthropic's new Opus flagship lands on OutcomeDev the day it ships. Same price as Opus 4.8, built for exactly the long-horizon agentic coding this platform runs on.
Pick it in the model selector and run it on your own Anthropic key. 1M token context, 128K max output.
Same $5 per million in, $25 per million out as Opus 4.8. A free upgrade in everything but name.
Fable 5 remains the tier above Opus for the hardest work and Sonnet 5 stays the recommended default. Opus 4.8, Opus 4.6, and Sonnet 4.6 are retired from the selector, superseded by Opus 5 and Sonnet 5 at the same price.
Moonshot AI's Kimi K3 lands on OutcomeDev the day it launches. The largest open-weight-bound model ever released, built for exactly what you do here: long-horizon agentic coding.
A new architecture generation with a 1M token context window and text, image, and video input. Weights promised open by July 27.
93.5% GPQA Diamond and 88.3% Terminal-Bench 2.1 at launch, the strongest open-weight results published, and #1 on the Frontend Code arena.
Pick K3 in the model selector and run it on your own Moonshot key. Sonnet-tier pricing at $3/$15 per million tokens.
The newest frontier models land on OutcomeDev the week they matter. Claude Fable 5 brings a new capability tier above Opus, Sonnet 5 becomes our recommended default, and Z.ai joins the multiverse with GLM-5.2.
Anthropic's most capable model, a new Mythos-class tier above Opus, built for the hardest long-horizon agentic work. 1M context.
Near-Opus quality on coding and agentic tasks at Sonnet speed. Now the recommended Claude model.
A new provider joins the model multiverse: Z.ai GLM-5.2 with a 1M token context window for whole-codebase navigation.
A week of dogfooding our own platform on camera surfaced and shipped a wave of improvements to the everyday workflow, from task creation to previews to the terminal.
Write "create a repo called invoicer and build..." and the platform creates the repository, clones it, and runs the task against it. Retries reuse the existing repo.
Sandboxes now expose the common dev-server ports and the preview follows wherever your app actually serves: Vite, Django, Flask, Angular, Expo web, and more.
The task workspace terminal now keeps your working directory across commands, explains itself on first use, and tells you how to wake a stopped sandbox.
New use cases for operators: AR aging, bank reconciliation, returns desks, lead routing with SLA timers, and more, plus persona hubs at /use-cases/for.
We have upgraded our model multiverse with the latest cutting-edge reasoning and coding models from Anthropic, MiniMax, and Moonshot AI.
Anthropic's latest flagship reasoning model is now available, delivering superior coherence and accuracy for complex engineering.
The new default MiniMax model featuring a massive 1M token context window and native high-speed token generation.
Moonshot AI's brand new "always-thinking" coding flagship. Replaces K2.6 with 30% fewer reasoning tokens and a high-speed engine.
Turn external alerts into actionable code. Our new public API allows your CI/CD, monitoring tools, or custom scripts to trigger AI tasks and real-time interjections.
Push notifications from GitHub Actions, Sentry, or any external service directly to your OutcomeDev dashboard.
Click one button to instantly pre-fill a Task Form with the error context, pre-selecting the repo and sandbox.
Inject alerts directly into active agent sessions. The agent will interrupt its current step and pivot to address the new context.
Issue and revoke developer tokens with SHA-256 hashing. Control your agent workforce programmatically and safely.
Moonshot AI’s 1T MoE model with 300 sub-agents is now available as a first-class coding agent on OutcomeDev.
300 specialized sub-agents orchestrate complex, long-horizon coding tasks with swarm intelligence.
Deep architectural reasoning across massive codebases — ideal for multi-file refactors and greenfield projects.
Competitive with Claude Opus on agentic benchmarks at a fraction of the cost. BYO key or use platform credits.
Two new frontier models from DeepSeek — a 1.6T flagship with 1M context and an ultra-efficient Flash variant that costs next to nothing.
The largest context window on OutcomeDev. Feed entire repositories into a single prompt without chunking.
The 284B Flash model delivers great quality at ~$0.14/M input tokens — the most cost-efficient option available.
V4 Pro matches Claude Opus on complex coding benchmarks while offering dramatically better cost efficiency.
Our most powerful iteration yet with the flagship Opus 4.7 model for unparalleled agentic reasoning.
Now supporting Anthropic’s flagship Opus 4.7 model for unparalleled agentic reasoning.
Deep reasoning and chain-of-thought capabilities for complex architectural decisions.
Best-in-class code generation, refactoring, and multi-file editing out of the box.
Expanding our agent workforce with MiniMax M2.5 and M2.7 high-performance MoE models.
Experience MiniMax models via our battle-tested Claude CLI infrastructure. Zero tool hallucination, maximum stability.
Deep architectural reasoning and massive file refactors are now standard for all coding tasks.
Internal specs and repository intents are now powered by a switchable, high-precision reasoning engine.
Completely automate any workflow with our new Scheduled Engine.
Trigger tasks "Every 9:00 AM" or using raw cron schedules natively.
OutcomeDev will checkout the repository, write code, run PRs, and clean up silently in the background.
Your jobs run reliably without ever overlapping using our new 3-tier task runtime.