Reupload
All AI ToolsLocal AI?
Back home

· Pandaily

StepFun releases Step 5 Preview LLM with 1M token context

StepFun releases Step 5 Preview LLM with 1M token context

Photo: Alina Grubnyak on Unsplash

StepFun announced Step 5 Preview on September 20, 2026, a 600-billion-parameter sparse Mixture-of-Experts model with 27 billion active parameters per token. The model features a one-million-token context window and scoring of 44 on Artificial Analysis' Intelligence Index, priced at $1 per million input tokens and $2.70 per million output tokens with a 95% cache discount.

StepFun officially launched Step 5 Preview on September 20, 2026, opening API access the same day. The model is a sparse Mixture-of-Experts architecture designed for long-horizon agent workloads including AI coding, software engineering, financial analysis, and professional knowledge work. It supports text and image inputs with a context window of one million tokens-enabling long documents and extensive code repositories to fit within a single API call.

The model employs a 92-layer narrow-deep Transformer layout rather than a wider network, with StepFun arguing that deeper architectures provide longer information paths for reasoning during extended prefill operations. To manage the computational demands of million-token sessions, Step 5 Preview incorporates Sparse Grouped-Query Attention with block-wise token merging, reducing indexer and top-k selection costs to approximately one-eighth of a denser baseline.

On the Artificial Analysis Intelligence Index, Step 5 Preview scores 44, matching Kimi K3 Max and placing it competitively while maintaining significantly lower costs than frontier models like GPT-5.6 Sol. The model generates output at 99.8 tokens per second with a time to first token of 2.96 seconds. Pricing is set at $1 per million input tokens and $2.70 per million output tokens, with a 95% cache discount available.

Open weights are scheduled for October 15, 2026, though the Hugging Face repository currently contains only a .gitattributes file. StepFun reports training emphasizing on-policy long-horizon reinforcement learning with techniques including load-aware scheduling, MTP-3 speculative decoding, FP8 MoE, and KV-cache offload, achieving more than threefold end-to-end acceleration for long-horizon reinforcement learning tasks.

Sources & credits

Original source: Pandaily