🔍 Read the full analysis: A Closer Look At Holo4’s Role In Computer-Use AI on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
H Company announced Holo4, a series of open-weight agentic models (27B dense and 35B-A3B MoE) that operate software through GUIs, code, MCP and APIs with a single model. The company reports 61.7% on OSWorld 2.0 for the 27B version, trailing only the strongest closed models, but all benchmark figures are self-reported.
H Company has released Holo4, a new series of open-weight agentic models designed to operate software through graphical interfaces, code, MCP and APIs using a single model, as detailed in the original analysis. The series comes in two sizes — a 27B dense model and a 35B-A3B Mixture of Experts model — both available through the H Models API and for download on Hugging Face in FP16, FP8 and GGUF formats. According to the company, the 27B version scores 61.7% on OSWorld 2.0, a computer-use benchmark, placing it behind only the strongest closed models, which it says it achieves at a fraction of the cost.
Holo4 is positioned as a generalist computer-use agent: it clicks and types on screens, writes and runs its own code, and calls MCP or API tools, selecting whichever interface fits the task. The same model runs on desktops, the web, Android, in a code sandbox and against business APIs, invoked the same way on each platform. H Company argues that most agentic models are trained for a single interface — GUI-focused models fail without a screen, while tool-calling models cannot handle applications that lack APIs.
The models were trained through supervised and reinforcement learning on a large set of environments and tasks, including tasks generated by H Company’s Agentic Task Factory. The company says Holo4 improves substantially over its Qwen base model, and it published side-by-side examples on professional software such as FreeCAD 3D modeling and Godot game design, run with the same prompt and harness. The dense model is built on Qwen3.8 27B and the MoE variant on Qwen3.6 35B-A3B, according to the company’s benchmark notes.
On benchmarks, H Company reports that on OSWorld 2.0 the 27B model scores 61.7% and the 35B-A3B model scores 30.9%, compared with 81.8% for Opus 5.5, which the company identifies as the strongest closed model. On AutomationBench for API use, Holo4 was measured in the company’s internal harness (v1.0.6) against public-set scores for other models. H Company has also open-sourced every trajectory behind its public benchmark scores, viewable at trajectories.hcompany.ai and downloadable from Hugging Face — an unusually transparent move at a time when AI firms debate how to prove content provenance.
Why Open-Weight Computer Agents Matter Now
The release matters because open-weight computer-use agents remain rare at this reported performance level. If Holo4’s scores hold up under independent evaluation, businesses could run capable software-automation agents at a fraction of the cost of frontier closed models, with the flexibility of self-hosting or open weights.
The cost-performance gap the company claims — 61.7% on OSWorld 2.0 from a 27B model against 81.8% from a much larger closed model — would represent meaningful progress for smaller, cheaper agents. The multi-interface design also addresses a practical limitation: real business tasks often mix screen work, code and API calls, and single-interface models break at those boundaries.
H Company’s decision to release all benchmark trajectories lets outside parties verify each step, which is more transparency than most closed-model providers offer.
From Holo1 to Holo4
Holo4 builds on H Company’s previous agentic models in the Holo series and arrives alongside an updated version of Holotron 3, called Holotron4 Nano. The new models are built on a Qwen base — Qwen3.8 27B for the dense model and Qwen3.6 35B-A3B for the MoE variant, according to the company’s benchmark notes.
H Company’s cost comparisons rest on specific assumptions: Holo4 is priced at H Models API rates for a single run, Qwen costs are calculated at Alibaba Cloud list prices (with cache hits at 20% of input price for the MoE model), and GPT and Opus effort sweeps come from OpenAI launch data. The company cautions that releases, harnesses and task subsets differ across the compared models.
“Real work is not siloed that way, and a single business task can require combining these different approaches.”
— H Company, announcement
Claims Awaiting Independent Verification
All headline benchmark numbers are self-reported by H Company and measured in the company’s own harness, which the company itself notes differs from other models’ releases, harnesses and task subsets.
On AutomationBench, other models’ scores come from the public set while costs come from a leaderboard running the private set — a mismatch the company acknowledges. Holo4 has not yet been evaluated on the AutomationBench private set.
The steep score difference between the 27B dense model (61.7%) and the larger 35B-A3B MoE model (30.9%) on OSWorld 2.0 is not explained in the announcement. Real-world reliability on business workflows, beyond curated demo examples, also remains unverified by third parties.
Evaluations and Adoption Watchpoints
H Company says it will report Holo4 results on the AutomationBench private set once that evaluation is complete. Independent benchmark submissions and third-party reproductions — now possible because trajectories and weights are public — will be the next test of the company’s claims.
Developers can access the models through the H Models API or download the full collection from Hugging Face in FP16, FP8 and GGUF formats.
Key Questions
What is Holo4?
Holo4 is a series of open-weight agentic models from H Company, released in two sizes — a 27B dense model and a 35B-A3B Mixture of Experts model — designed to operate software through GUIs, code, MCP and APIs using a single model.
How does Holo4 perform on benchmarks?
According to H Company, the 27B model scores 61.7% on OSWorld 2.0 and the 35B-A3B model scores 30.9%, compared with 81.8% for Opus 5.5. All of these figures are self-reported by the company and measured in its own harness.
Where can developers get Holo4?
The models are available through the H Models API and can be downloaded from Hugging Face in FP16, FP8 and GGUF formats.
Why does the 35B MoE model score lower than the smaller 27B model?
The announcement does not explain the roughly 30-point gap on OSWorld 2.0 between the two model sizes. This remains one of several unanswered questions about the release.
Have independent evaluators confirmed Holo4’s results?
No. All headline numbers are self-reported, though H Company has published every benchmark trajectory, allowing outside parties to audit and attempt to reproduce the results.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
