AI News Feed
Market watch
Large Language Models

H Company Releases Holo4 Agentic Models in Two Sizes, Adds Holotron4 Nano

H Company launched Holo4, a new agentic model series in 27B dense and 35B-A3B MoE sizes, plus Holotron4 Nano.

The post said Holo4 builds on the previous model and interacts with software through any available interface, including graphical user interfaces, code, MCP and APIs. It clicks and types on a screen, writes and runs its own code, and calls MCP or API tools, using whichever fits the task.

Holo4 was trained through supervised and reinforcement learning on a large set of environments and tasks, including those generated by the company's Agentic Task Factory. The company said the models score well on academic benchmarks but are built for real business workflows.

The post said most agentic models are trained for one interface only. GUI-focused models are blind without a screen, while models that prefer tool calling are stuck in front of an application that has no API. H Company said real work is not siloed that way, and a single business task can require combining these approaches. Holo4 runs on desktops, on the web, on Android, in a code sandbox and against business APIs, with the same model called the same way in each case, so users do not need to select a different model for each platform.

On OSWorld 2.0, Holo4 27B scores 61.7 percent, against 81.8 percent for Opus 5.5, while Holo4 35B-A3B reaches 30.9 percent, according to the post. H Company said Holo4 trails only the strongest closed models on long workflows but does so with orders of magnitude fewer parameters and at much lower cost. The company also said the models improve significantly over their Qwen base. On the hardest academic benchmarks for desktop control, OSWorld 2.0, and API use, AutomationBench, Holo4 competes with frontier models at a much lower cost per task, the post said.

H Company said it open-sources every trajectory behind its scores on public benchmarks. Each step can be replayed at trajectories.hcompany.ai or downloaded from Hugging Face. The post includes links to Holo4-27B, Holo4-35B-A3B, Holotron4 Nano, a full collection in FP16, FP8 and GGUF formats, a trajectory viewer and dataset, and an H Models API quickstart.

For OSWorld 2.0 cost-performance estimates, the company said costs are estimated from the input and output tokens of each agentic run. Holo4 is priced at H Models API rates for a single run. Qwen3.8 27B uses a model card score, with cost from the tokens of the company's run at Alibaba Cloud list prices. Qwen3.6 35B-A3B uses a single run in the company's harness at Alibaba Cloud list prices with cache hits at 20 percent of the input price. OpenAI launch data supplies the GPT and Opus effort sweeps, while other closed and open-weight points use the official OSWorld 2.0 leaderboard. Releases, harnesses and task subsets differ, and the line connects non-dominated score and cost pairs among the closed models, with Holo4 excluded, according to the post.

For AutomationBench, Holo4, Qwen3.8 27B and Qwen3.6 35B-A3B were measured on AutomationBench v1.0.6, with scores and costs measured in H Company's internal harness. Other models use public-set scores from the AutomationBench README and cost per task from the official leaderboard, which runs on the private set. H Company said it will report Holo4 on the private set once it is evaluated.

The post also showed Holo4 27B alongside Qwen3.8 27B, its base model, on professional software tasks with the same prompt and harness for both models. In one example, building a 3D model of the Eiffel Tower in FreeCAD to a detailed design, Holo4 27B used 84 calls and 1.3 million tokens, while Qwen3.8 27B used 60 calls and 1.0 million tokens.