Cohere Launches North Mini Code, its First Open-Source Agentic Coding Model

On June 9, 2026, Cohere announced the release of North Mini Code, its first open-source agentic coding model. The Mixture-of-Experts (MoE) model features 30 billion total parameters, with 3 billion active during generation, and is designed to deliver high-performance software eng

Cohere Launches North Mini Code, its First Open-Source Agentic Coding Model
Cohere Launches North Mini Code, its First Open-Source Agentic Coding Model

On June 9, 2026, Cohere announced the release of North Mini Code, its first open-source agentic coding model. The Mixture-of-Experts (MoE) model features 30 billion total parameters, with 3 billion active during generation, and is designed to deliver high-performance software engineering capabilities without requiring extensive hardware. Released under an Apache 2.0 license, the model supports a total context length of 256,000 tokens and a maximum generation length of 64,000 tokens.

Model Specifications and Hardware Requirements

North Mini Code (version 1.0) is optimized for code generation, agentic software engineering, and terminal tasks. The model is distributed in multiple precision formats, including bf16, fp8, and w4a16.

According to Cohere, running the model locally or on-premises requires a single Nvidia H100 GPU running at FP8 precision, or a single Nvidia H100 GPU running at FP4 precision.

Agentic Capabilities and Benchmark Performance

Cohere designed North Mini Code specifically for agentic workflows. This includes tasks such as coordinating and understanding sub-agents, mapping out systems architecture, and conducting automated code reviews.

The company evaluated the model using specialized software engineering harnesses. These evaluations utilized the “SWE-agent” harness for both the SWE-Bench Verified and SWE-Bench Pro benchmarks. Evaluations also spanned Terminal-Bench v2 and Terminal-Bench Hard.

Based on these tests, North Mini Code achieved a score of 33.4 on the Artificial Analysis Coding Index.

Throughput and Latency Measurements

In internal testing comparing the model to Devstral Small 2 under identical concurrency levels and hardware configurations, Cohere reported the following performance metrics:

  • Output Throughput: North Mini Code achieved up to 2.8x higher output throughput than Devstral Small 2, allowing for a faster work rate under heavy workloads.
  • Inter-Token Latency: The model delivered a 30% lower inter-token latency compared to Devstral Small 2, indicating more consistent token-generation pacing.
  • Time-to-First-Token (TTFT): In contrast to the throughput advantages, Devstral Small 2 maintained a slight performance edge in prompt processing speed (TTFT), with North Mini Code running approximately 5% slower under the tested conditions.

Deployment and Platform Availability

The model is integrated with the open-source OpenCode tool, though it is compatible with most standard coding agents. Developers can access and run North Mini Code through several distribution channels:

  • Hugging Face: Model weights are openly downloadable in bf16, fp8, and w4a16 formats.
  • Cohere API & OpenRouter: Available for direct integration and querying.
  • Model Vault: Cohere’s managed enterprise inference platform offers dedicated deployment options.
Topics
  • #Startups
Krishnan

Author

Krishnan

Contributor

Enterprise Technology Explorer is a business and operations professional with over 15 years of experience across multiple industries working with Fortune 500 companies. With a solid foundation in enterprise processes, digital adoption, and technology evaluation, he excels at bridging business needs with emerging technologies to build scalable enterprise-grade applications.