AIARTICLE

H Company releases Holo4, an open model for agents that operate screen, code, and MCP

Released on September 28, 2026, the model unifies GUI control, code execution, and API calls into a single open-source weight, but it still falls short of closed models on long tasks.

H Company published the weights for Holo4, its new generation of agentic models for "computer use," on the Hugging Face blog on September 28, 2026: agents that click, type, write code, and call tools to operate software end to end. The announcement matters to anyone building agents because it addresses a practical architecture problem: until now, the choice was between models strong at GUI (which are blind without a screen) or models strong at tool calling (which stall in front of an app with no API). Holo4 is marketed as a single model that decides on its own which interface to use at each step of the task.

One model, four ways to act

Holo4 comes in two sizes: a dense 27B-parameter model and a Mixture of Experts with 35B total parameters and 3B active per token (35B-A3B). According to H Company, the same checkpoint works on desktop, web, Android, code sandboxes, and against real business APIs, without swapping models per platform. This matters for anyone currently building agent pipelines with a vision model for the GUI part and a separate model for tool calling: the pitch is to collapse those two pieces into one, which simplifies the harness and reduces maintenance cost.

The models are built on the Qwen base (Qwen3.8 27B and Qwen3.6 35B-A3B) and were trained with supervised learning and reinforcement learning over a large set of environments generated by the company's own "Agentic Task Factory," an internal pipeline that creates interactive environments and verifiable tasks from documentation, screenshots of real websites, and open-source software. The company says it produced about 10,000 tasks this way, including hybrid environments that expose the same state through both GUI and MCP.

The numbers: where it gets close and where it still falls behind

On OSWorld 2.0, the standard benchmark for desktop control on long tasks, Holo4 27B scored 61.7%, against 81.8% for Opus 5.5. Holo4 35B-A3B, smaller in active parameters, came in at 30.9%. H Company acknowledges the gap and frames its argument around cost per task: in the cost-benefit charts it published, Holo4 shows up competing with closed models at a fraction of the price, since cost is estimated from the input and output tokens of each run, billed at the H Models API's own rates.

Table comparing Holo4 with closed models such as Opus 5.5 and GPT-5.6 Sol on OSWorld 2.0, showing score, success rate, and cost per task
Table comparing Holo4 with closed models such as Opus 5.5 and GPT-5.6 Sol on OSWorld 2.0, showing score, success rate, and cost per task. Reproduction: huggingface.co.

It's worth noting the caveat the company itself makes in the fine print of the charts: releases, harnesses, and task subsets vary across the benchmarks being compared, so the direct comparison with Opus 5.5 or GPT-5.6 Sol carries a margin of error. For AutomationBench (which measures API usage), the pattern is similar: H Company runs its own internal harness for Holo4 and for the base Qwen models, but uses official leaderboard numbers for the closed competitors, which run on a different, private dataset. It's worth checking that detail before citing the numbers as an unbiased comparison.

A more concrete and verifiable data point comes from the usage examples the company published: in a task to generate a Pac-Man-style game in Godot, Holo4 27B completed the task in 68 agent calls and 2.4 million tokens, versus 197 calls and 11.4 million tokens for Qwen3.8 27B (the base model, without the agentic post-training). In another task, modeling the H Company logo in FreeCAD, Holo4 used 94 calls and 1.5 million tokens, versus 118 calls and 1.9 million tokens for the base Qwen. The pattern repeats across the published examples: Holo4 tends to solve the same task with fewer iterations and fewer tokens spent, which translates directly into lower execution cost, though comparing the final quality of the generated artifact requires visually inspecting the results, available in the original post.

How to run it: open weights, but not trivially plug-and-play

Here's the point that matters to anyone looking to move away from closed-API dependency. Holo4's weights are published on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF, alongside the smaller Holotron4 Nano model. The existence of a quantized GGUF is what makes it feasible to run a version of the model on developer hardware, although the dense 27B Holo4 in FP16 still requires a GPU with plenty of VRAM, and the 35B-A3B MoE has a larger model memory footprint even with few active parameters per token. Anyone without that hardware available can use the H Models API, which the company also offers in both sizes since launch.

An architectural detail that changes the risk calculus for anyone putting this into production: H Company rebuilt the harness (the loop that executes the model's actions and manages its context over hundreds of steps) to give the agent reliable memory over long sequences and, more sensitively, a shell on the desktop machine itself. That means the agent isn't limited to clicking on the screen: it can open a terminal and run commands directly. For anyone deploying Holo4 locally, this raises the obvious question of sandboxing and permissions, something H Company's material doesn't detail and that falls on whoever integrates the model into a real pipeline.

Holotron4 Nano and the bet on the Nemotron coalition

Alongside Holo4, the company also released Holotron4 Nano, an update to Holotron 3, applying the same agentic post-training recipe to NVIDIA's Nemotron 3 Nano Omni model. H Company is part of the Nemotron Coalition, an NVIDIA initiative to expand the Nemotron model ecosystem, and it's using this release as a proof of concept that the training recipe (the Agentic Task Factory plus the rebuilt harness) generalizes to other base architectures, and isn't tied to the size or the Qwen family used in the main Holo4. For anyone already running Nemotron in production for other reasons, this opens a path to add agentic capability without switching model families.

When it's not worth switching

If the workflow in question is a long, complex desktop automation task, where success rate matters more than cost per run, the gap to Opus 5.5 (81.8% versus 61.7% on OSWorld 2.0) is still wide enough to justify not abandoning the closed model. Holo4 makes more sense in high-volume scenarios and more constrained tasks, where cost per call weighs on the budget and the 20-percentage-point accuracy gap can be offset by retries or by a human in the loop reviewing the result. For anyone who just wants to test without committing infrastructure, the H Models API removes the need for hardware, but it also removes part of the independence appeal that motivates looking at an open-source model in the first place.

Translated from the Brazilian Portuguese original · Read the original

View profile →