2026 itvalue文章顶部

27B Local Model From StartLux Ranks Second in Official Chinese Agent Benchmark, Matching Larger Rivals on Key Tasks

A 27-billion-parameter local model from Shanghai-based StartLux scored 39.25 and ranked second in a China Academy of Information and Communications Technology agent-capability test, according to multiple technology outlets citing the academy’s report. The model outperformed larger DeepSeek variants on several multi-tool tasks.

NextFin News — A compact local model has recorded a strong result in an official Chinese evaluation of agent performance. StartLux-V1.0-27B-Preview, developed by Shanghai Yuandian Xinghui Science and Technology (StartLux), achieved a composite score of 39.25 and placed second overall in the Trusted AI Large Model Benchmark’s MCP special test, according to multiple technology outlets summarizing a report from the China Academy of Information and Communications Technology’s artificial-intelligence institute.

The evaluation covered six specialized tasks—location navigation, web search, browser automation, financial analysis, code-repository management and 3D design—plus an overall assessment. It examined multi-tool coordination, complex task execution and interaction with realistic environments, drawing in part on the open-source MCP-Universe framework. Models listed in the reported comparison included DeepSeek-V4-Pro (cited as 1.6 trillion parameters), DeepSeek-V4-Flash-0731 (cited as 284 billion), Step-3.7-Flash (cited as 198 billion), the same-size Qwen-3.6-27B, and a 4-billion-parameter entry identified as AgentCPM-Explore.

According to the secondary accounts, StartLux’s 27-billion-parameter system finished ahead of the 284-billion-parameter DeepSeek Flash model and the 198-billion-parameter Step model. Against Qwen-3.6-27B it recorded a 5.34-point advantage. In individual categories it ranked first in location navigation and tied or led the much larger DeepSeek-V4-Pro in browser automation and financial analysis.

StartLux states that the model began from a Qwen3.6-27B base and received targeted post-training. The company describes the process as relying on an “AI-trains-AI” or Auto Research loop in which the system designs experiments, evaluates results and iteratively refines strategy. It presents the 27B preview as the first domestic local agent model trained in this manner.

The model is designed to run on consumer-grade personal computers. StartLux has said it plans to release a first-generation local intelligent solution later this year and continues additional training work, including exploration of alternative architectures.

The reported ranking arrives amid rising interest in models that keep data and inference on-device. Google, Meta and NVIDIA have each released open or open-weight models in a roughly 30-billion-parameter range intended for local or edge deployment. Once base capability clears a practical threshold, many users and organizations place higher value on data locality, controllable cost and the ability to customize than on pure scale.

Agent-style tests of the kind used in the CAICT MCP evaluation are particularly relevant to this shift. Simple conversational accuracy no longer captures the full requirement. Practical usefulness increasingly depends on the ability to call tools, maintain state across steps and operate in changing environments. A compact model that performs well on such tasks can, in principle, deliver value without the infrastructure demands of far larger systems.

Whether the reported scores will hold under independent reproduction and wider testing remains open. Single-benchmark results are always provisional. Real-world latency on consumer hardware, robustness over longer sessions, and the quality of the surrounding software stack will matter at least as much as the headline ranking. Still, the figures summarized from the CAICT report supply a concrete data point: a 27-billion-parameter local model has matched or exceeded much larger systems on a set of multi-tool agent tasks under an official evaluation framework.

StartLux frames local deployment not as a compromise but as the setting in which privacy, cost control and customization are most cleanly addressed. The next test will be whether laboratory rankings translate into reliable everyday performance and whether the planned local solution attracts developers and organizations that have so far relied on cloud services.

The episode underscores a quiet rebalancing in the field. Scale continues to matter, yet it is no longer the sole determinant of usefulness. Models that keep computation close to the data and the user are beginning to appear competitive on the tasks that matter most to practical agents. If sustained, that change would reshape both the economics of deployment and the kinds of systems organizations choose to run.

本文系作者 Chelsea_Sun 授权钛媒体发表,并经钛媒体编辑,转载请注明出处、作者和本文链接
本内容来源于钛媒体钛度号,文章内容仅供参考、交流、学习,不构成投资建议。
想和千万钛媒体用户分享你的新奇观点和发现,点击这里投稿 。创业或融资寻求报道,点击这里

敬原创,有钛度,得赞赏

赞赏支持
发表评论
0 / 300

根据《网络安全法》实名制要求,请绑定手机号后发表评论

登录后输入评论内容

扫描下载App