What's New

The complete history of new features, changes, and upgrades shipped to help you become a better Outcome Developer.

Latest ReleaseJuly 24, 2026

Claude Opus 5 Is Here, Day One

Anthropic's new Opus flagship lands on OutcomeDev the day it ships. Same price as Opus 4.8, built for exactly the long-horizon agentic coding this platform runs on.

Opus 5, available now

Pick it in the model selector and run it on your own Anthropic key. 1M token context, 128K max output.

Flagship power, unchanged pricing

Same $5 per million in, $25 per million out as Opus 4.8. A free upgrade in everything but name.

The Claude lineup, refreshed

Fable 5 remains the tier above Opus for the hardest work and Sonnet 5 stays the recommended default. Opus 4.8, Opus 4.6, and Sonnet 4.6 are retired from the selector, superseded by Opus 5 and Sonnet 5 at the same price.

Try Opus 5
Feature ReleaseJuly 16, 2026
Kimi

Kimi K3: The 2.8 Trillion Parameter Flagship, Day One

Moonshot AI's Kimi K3 lands on OutcomeDev the day it launches. The largest open-weight-bound model ever released, built for exactly what you do here: long-horizon agentic coding.

2.8T MoE with Kimi Delta Attention

A new architecture generation with a 1M token context window and text, image, and video input. Weights promised open by July 27.

Frontier agentic benchmarks

93.5% GPQA Diamond and 88.3% Terminal-Bench 2.1 at launch, the strongest open-weight results published, and #1 on the Frontend Code arena.

Available now, BYOK

Pick K3 in the model selector and run it on your own Moonshot key. Sonnet-tier pricing at $3/$15 per million tokens.

Try Kimi K3
Feature ReleaseJuly 10, 2026

The Claude 5 Era: Fable 5 and Sonnet 5 Arrive (plus Z.ai GLM-5.2)

The newest frontier models land on OutcomeDev the week they matter. Claude Fable 5 brings a new capability tier above Opus, Sonnet 5 becomes our recommended default, and Z.ai joins the multiverse with GLM-5.2.

Claude Fable 5

Anthropic's most capable model, a new Mythos-class tier above Opus, built for the hardest long-horizon agentic work. 1M context.

Claude Sonnet 5 (new default)

Near-Opus quality on coding and agentic tasks at Sonnet speed. Now the recommended Claude model.

Z.ai GLM-5.2 with 1M context

A new provider joins the model multiverse: Z.ai GLM-5.2 with a 1M token context window for whole-codebase navigation.

Try the new models
Feature ReleaseJuly 10, 2026

Platform Week: Plain-Language Repos, Universal Previews, and a Real Terminal

A week of dogfooding our own platform on camera surfaced and shipped a wave of improvements to the everyday workflow, from task creation to previews to the terminal.

Create repositories with plain language

Write "create a repo called invoicer and build..." and the platform creates the repository, clones it, and runs the task against it. Retries reuse the existing repo.

Previews for every framework

Sandboxes now expose the common dev-server ports and the preview follows wherever your app actually serves: Vite, Django, Flask, Angular, Expo web, and more.

A terminal that behaves like one

The task workspace terminal now keeps your working directory across commands, explains itself on first use, and tells you how to wake a stopped sandbox.

20 Digital Operations blueprints

New use cases for operators: AR aging, bank reconciliation, returns desks, lead routing with SLA timers, and more, plus persona hubs at /use-cases/for.

See the use cases
Feature ReleaseJune 15, 2026

June 2026 Frontier Models: Opus 4.8, MiniMax M3, & Kimi K2.7-Code

We have upgraded our model multiverse with the latest cutting-edge reasoning and coding models from Anthropic, MiniMax, and Moonshot AI.

Claude Opus 4.8

Anthropic's latest flagship reasoning model is now available, delivering superior coherence and accuracy for complex engineering.

MiniMax M3 MoE

The new default MiniMax model featuring a massive 1M token context window and native high-speed token generation.

Kimi K2.7-Code

Moonshot AI's brand new "always-thinking" coding flagship. Replaces K2.6 with 30% fewer reasoning tokens and a high-speed engine.

Explore Models
Feature ReleaseMay 11, 2026

Actionable Developer Notifications API

Turn external alerts into actionable code. Our new public API allows your CI/CD, monitoring tools, or custom scripts to trigger AI tasks and real-time interjections.

Public Notify Endpoint

Push notifications from GitHub Actions, Sentry, or any external service directly to your OutcomeDev dashboard.

Actionable 'Start Task'

Click one button to instantly pre-fill a Task Form with the error context, pre-selecting the repo and sandbox.

Real-time Interjection

Inject alerts directly into active agent sessions. The agent will interrupt its current step and pivot to address the new context.

Secure Token Management

Issue and revoke developer tokens with SHA-256 hashing. Control your agent workforce programmatically and safely.

View Developer Settings
Feature ReleaseMay 5, 2026
Kimi

Kimi K2.6 — Agent Swarm Architecture

Moonshot AI’s 1T MoE model with 300 sub-agents is now available as a first-class coding agent on OutcomeDev.

1 Trillion Parameter MoE

300 specialized sub-agents orchestrate complex, long-horizon coding tasks with swarm intelligence.

256K Context Window

Deep architectural reasoning across massive codebases — ideal for multi-file refactors and greenfield projects.

Top SWE-Bench Performer

Competitive with Claude Opus on agentic benchmarks at a fraction of the cost. BYO key or use platform credits.

Try Kimi Agent
Feature ReleaseMay 5, 2026
DeepSeek

DeepSeek V4 Pro & Flash

Two new frontier models from DeepSeek — a 1.6T flagship with 1M context and an ultra-efficient Flash variant that costs next to nothing.

1 Million Token Context

The largest context window on OutcomeDev. Feed entire repositories into a single prompt without chunking.

V4 Flash — 20x Cheaper

The 284B Flash model delivers great quality at ~$0.14/M input tokens — the most cost-efficient option available.

Opus-Tier Reasoning

V4 Pro matches Claude Opus on complex coding benchmarks while offering dramatically better cost efficiency.

Try DeepSeek Agent
Feature ReleaseApril 21, 2026

Claude Opus 4.7

Our most powerful iteration yet with the flagship Opus 4.7 model for unparalleled agentic reasoning.

Claude Opus 4.7 Flagship

Now supporting Anthropic’s flagship Opus 4.7 model for unparalleled agentic reasoning.

Extended Thinking

Deep reasoning and chain-of-thought capabilities for complex architectural decisions.

Production-Grade Coding

Best-in-class code generation, refactoring, and multi-file editing out of the box.

Start New Task
Feature ReleaseMarch 29, 2026
Minimax

Native MiniMax AI Integration

Expanding our agent workforce with MiniMax M2.5 and M2.7 high-performance MoE models.

Industrial Proxy Architecture

Experience MiniMax models via our battle-tested Claude CLI infrastructure. Zero tool hallucination, maximum stability.

200k Context Window

Deep architectural reasoning and massive file refactors are now standard for all coding tasks.

Unified Taskmaster Utility

Internal specs and repository intents are now powered by a switchable, high-precision reasoning engine.

Try MiniMax Agent
Feature ReleaseMarch 27, 2026

Introducing Scheduled Tasks

Completely automate any workflow with our new Scheduled Engine.

Cron Expression Engine

Trigger tasks "Every 9:00 AM" or using raw cron schedules natively.

Autonomous Execution

OutcomeDev will checkout the repository, write code, run PRs, and clean up silently in the background.

Idempotent Trigger Routing

Your jobs run reliably without ever overlapping using our new 3-tier task runtime.

View Dashboard