On September 1st, iFLYTEK's wholly-owned subsidiary, Ciyuan Xinghuo, officially launched and open-sourced two edge-side general large models: Xinghuo X2.5-4B and Xinghuo X2.5-1.7B. These two products are the first edge-side models to natively support a context length of up to 1 million Tokens, and have fully opened model weights, code repositories, and deployment documentation.
Million-level Context Support, Small Models Can "Read the Whole Book"
Both models use a hybrid attention architecture, optimized for core capabilities such as agents, code, mathematics, and instruction following, and their comprehensive tested performance is leading among industry-class open-source models of similar size.
Regarding context capability, Xinghuo X2.5-4B and Xinghuo X2.5-1.7B are trained on high-quality data at the level of trillions of Tokens, covering scenarios such as long documents and technical materials. They natively support a context window of 1 million Tokens, allowing them to receive and understand larger-scale information at once.

For example, in product after-sales service: users can import the complete after-sales manual into Xinghuo X2.5-4B. The model first sorts out after-sales rules under different scenarios and marks corresponding chapters. When asked "Whether a device failure within 10 days of purchase and using third-party consumables qualifies for a replacement?", it can provide a comprehensive judgment by linking regulations on returns, fault handling, and exceptions across chapters. Even if new conditions such as remote areas or devices with family maps are added later, the model can maintain the previous context and continue to provide advice by combining logistics, data erasure, and cost-bearing rules.
The long context solves the issue of "whether information can be fully viewed," while the agent and tool calling capabilities solve "whether actions can be taken after viewing." Xinghuo X2.5 edge-side models provide localized intelligent capabilities for scenarios such as personal office, code development, smart hardware, and robotics. In code development scenarios, Xinghuo X2.5-4B can match cloud models with parameters 2 to 3 times larger in tasks such as algorithm implementation, code completion, and generation; on the Domux smart home test set, the end-to-end execution accuracy of Xinghuo X2.5-1.7B for control instructions reaches 90.3%, with an average response time of only 0.85 seconds; in robotics scenarios, both models can be deployed on robot bodies or edge devices, supporting tasks such as operation control, target tracking, and navigation decision-making, reducing dependency on cloud connections and fixed pre-set programs.
Entirely Domestic Computing Power Training, Instant Download and Experience
Xinghuo X2.5-4B and Xinghuo X2.5-1.7B were trained entirely on domestic computing platforms, using about 20 trillion diverse tokens of data for pre-training, and continuously improving performance through high-quality supervised fine-tuning and reinforcement learning.
In terms of deployment, both models support hardware platforms including NVIDIA, Huawei, Hai Guang, and Hemo, and are compatible with mainstream inference frameworks such as vLLM, SGLang, and llama.cpp. They can be quickly deployed through tools like Ollama and LM Studio, and support incremental training using LLaMA-Factory. Starting from today, model weights have been made available on platforms such as Hugging Face and GitHub, and corresponding model APIs have been launched on iFLYTEK's Starry MaaS platform, free of charge for a limited time. On September 7th, iFLYTEK will officially release the next-generation flagship general large model, Xinghuo X2.5, further upgrading core capabilities such as code and agents.



