NextFin News — A compact local model has recorded a strong result in an official Chinese evaluation of agent performance. StartLux-V1.0-27B-Preview, developed by Shanghai Yuandian Xinghui Science and Technology (StartLux), achieved a composite score of 39.25 and placed second overall in the Trusted AI Large Model Benchmark’s MCP special test, according to multiple technology outlets summarizing a report from the China Academy of Information and Communications Technology’s artificial-intelligence institute.
The evaluation covered six specialized tasks—location navigation, web search, browser automation, financial analysis, code-repository management and 3D design—plus an overall assessment. It examined multi-tool coordination, complex task execution and interaction with realistic environments, drawing in part on the open-source MCP-Universe framework. Models listed in the reported comparison included DeepSeek-V4-Pro (cited as 1.6 trillion parameters), DeepSeek-V4-Flash-0731 (cited as 284 billion), Step-3.7-Flash (cited as 198 billion), the same-size Qwen-3.6-27B, and a 4-billion-parameter entry identified as AgentCPM-Explore.
According to the secondary accounts, StartLux’s 27-billion-parameter system finished ahead of the 284-billion-parameter DeepSeek Flash model and the 198-billion-parameter Step model. Against Qwen-3.6-27B it recorded a 5.34-point advantage. In individual categories it ranked first in location navigation and tied or led the much larger DeepSeek-V4-Pro in browser automation and financial analysis.
StartLux states that the model began from a Qwen3.6-27B base and received targeted post-training. The company describes the process as relying on an “AI-trains-AI” or Auto Research loop in which the system designs experiments, evaluates results and iteratively refines strategy. It presents the 27B preview as the first domestic local agent model trained in this manner.
The model is designed to run on consumer-grade personal computers. StartLux has said it plans to release a first-generation local intelligent solution later this year and continues additional training work, including exploration of alternative architectures.
The reported ranking arrives amid rising interest in models that keep data and inference on-device. Google, Meta and NVIDIA have each released open or open-weight models in a roughly 30-billion-parameter range intended for local or edge deployment. Once base capability clears a practical threshold, many users and organizations place higher value on data locality, controllable cost and the ability to customize than on pure scale.
Agent-style tests of the kind used in the CAICT MCP evaluation are particularly relevant to this shift. Simple conversational accuracy no longer captures the full requirement. Practical usefulness increasingly depends on the ability to call tools, maintain state across steps and operate in changing environments. A compact model that performs well on such tasks can, in principle, deliver value without the infrastructure demands of far larger systems.
Whether the reported scores will hold under independent reproduction and wider testing remains open. Single-benchmark results are always provisional. Real-world latency on consumer hardware, robustness over longer sessions, and the quality of the surrounding software stack will matter at least as much as the headline ranking. Still, the figures summarized from the CAICT report supply a concrete data point: a 27-billion-parameter local model has matched or exceeded much larger systems on a set of multi-tool agent tasks under an official evaluation framework.
StartLux frames local deployment not as a compromise but as the setting in which privacy, cost control and customization are most cleanly addressed. The next test will be whether laboratory rankings translate into reliable everyday performance and whether the planned local solution attracts developers and organizations that have so far relied on cloud services.
The episode underscores a quiet rebalancing in the field. Scale continues to matter, yet it is no longer the sole determinant of usefulness. Models that keep computation close to the data and the user are beginning to appear competitive on the tasks that matter most to practical agents. If sustained, that change would reshape both the economics of deployment and the kinds of systems organizations choose to run.







快报
根据《网络安全法》实名制要求,请绑定手机号后发表评论