Nex-AGI Drops Apache-2.0 Weights for Nex-N2.5 Mini Computer-Use Model

Nex-AGI published Apache-2.0 weights for Nex-N2.5 Mini (~35B/~3B active multimodal) on Hugging Face September 8, 2026, with a 2×H100 SGLang recipe; Pro multimodal weights remain coming soon.

Nex-N2.5 Mini Apache-2.0 computer-use open weights
Nex-N2.5 Mini Apache-2.0 computer-use open weights

Nex-AGI introduced its Nex-N2.5 agentic family on September 8, 2026, and the downloadable tier is Nex-N2.5 Mini: Apache-2.0 multimodal weights on Hugging Face (nex-agi/Nex-N2.5-mini), roughly 35 billion total parameters with about 3 billion active per token, sixteen BF16 safetensors shards totaling about 70 GB, and a documented two-H100 reference deployment. OrcaRouter’s same-day unpack frames the effort as a collaboration across the Shanghai Innovation Institute, Shanghai Qiji Zhifeng, Mosi Intelligence, and Kuafu Technology. The larger multimodal sibling Nex-N2.5 Pro is still marked weights “coming soon,” while text-only Nex-N2.5 Max (1.6T / ~49B active) is also up but targets sixteen H200s across two nodes.

Mini continues the multimodal Nex-N2 line for computer use, browsing, and visually grounded agent work. The released config is a Qwen3_5MoeForConditionalGeneration / qwen3_5_moe stack: 40 layers, hidden size 2048, grouped-query attention with 16 query heads and 2 KV heads, vocabulary 248,320, MoE with 256 routed experts and 8 active plus a shared expert, and max_position_embeddings of 262,144. The Hugging Face card lists the model as ungated BF16 and not behind a gate; downloads were still minimal on day one.

Serving recipe and thinking modes

Nex-AGI ships a prebuilt SGLang image, nexagi/sglang:v0.5.18-nex-patch. For Mini the documented launch is a single node with two H100s and tensor parallelism 2, --reasoning-parser qwen3, and --tool-call-parser qwen3_coder. Recommended sampling is temperature 0.7, top_p 0.95, and top_k 40. Thinking is controlled with reasoning_effort: "none" (no reasoning trace), "medium" (adaptive default), or "high" (always think). Pro’s recipe is eight H100s on the same image once shards appear; Max uses a different DeepSeek-based parser path and multi-node EP/TP settings.

Vendor benchmarks — label them as Nex-AGI’s

The model card’s scores are vendor-reported via NexAU (coding) and NexCUA (computer/browser use), with NexCUA slated to be open-sourced. OrcaRouter stresses that no independent reproduction was published at writing, so treat the rows as Nex-AGI’s case until a neutral run lands. Mini’s listed figures include Terminal-Bench 2.1 at 73.4, OSWorld-Verified at 71.2, WebArena-Verified at 63.4, BrowseComp at 83.4, SWE-Bench Pro at 43.8, Toolathlon Verified at 54.6, OmniDoc at 89.7, and a harsher OSWorld-2 at 30.5—a different suite from OSWorld-Verified, not a contradiction of the verified score.

For teams evaluating open computer-use agents, Mini is the on-ramp: permissive license, two-H100 footprint, and a published tool-call plus reasoning-effort serving path. Hold for Pro if the priority is the family’s strongest multimodal computer-use tier—those weights were not downloadable at launch.

Primary sources for this brief are OrcaRouter’s Nex-N2.5 Mini open-weights unpack (https://www.orcarouter.ai/blog/nex-n2-5-mini-open-weights-release) and the Hugging Face nex-agi/Nex-N2.5-mini model card (https://huggingface.co/nex-agi/Nex-N2.5-mini).

Topics
  • #Opensource
  • #AI Agents
  • #Products
Raj M

Author

Raj M

Contributor

AI Systems Architect is a seasoned technology leader with over 15 years of experience in the IT industry working with Fortune 500 companies. With a solid foundation in multi-agent systems, open-source LLM infrastructure, and enterprise deployment, he excels at building scalable production-grade AI platforms.