Exploring Holo4 And The Future Of Computer-Use AI
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Exploring Holo4 And The Future Of Computer-Use AI on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

H Company has released two open-weight Holo4 models designed to use graphical interfaces, code, MCP and APIs within one system. The company reports a 61.7% OSWorld 2.0 score for its 27B model, but the results have not been independently verified and comparisons use different evaluation setups.

H Company has released Holo4, a series of open-weight models built to operate software through graphical interfaces, code, MCP and APIs. The release includes a 27B dense model and a 35B-A3B Mixture of Experts model; H Company says the 27B version scored 61.7% on OSWorld 2.0, a benchmark for computer-use tasks.

The models are available through the H Models API and for download from Hugging Face in FP16, FP8 and GGUF formats. H Company describes Holo4 as a general-purpose computer-use agent: it can click and type on a screen, write and run code, or call MCP and API tools, depending on the task. The company says the same model can be used across desktops, the web, Android, a code sandbox and business APIs.

H Company reports that Holo4 27B scored 61.7% on OSWorld 2.0, while Holo4 35B-A3B scored 30.9%. The company compares those figures with 81.8% for Opus 5.5, which it identifies as the strongest closed model in its comparison. It says the 27B result comes at a fraction of the cost, though its comparisons rely on differing releases, harnesses and task subsets.

The company says the models were trained with supervised and reinforcement learning on tasks from a range of environments, including tasks generated by its Agentic Task Factory. It also points to side-by-side examples in FreeCAD and Godot, run with the same prompt and harness, as evidence of improvement over the Qwen base models. Those examples and performance claims are company-reported; the source material does not cite independent evaluations.

At a glance
announcementWhen: Announced; the source material does not…
The developmentH Company released the Holo4 series, a pair of open-weight models designed to handle software tasks across graphical interfaces, code and tools.
At a glance
announcementWhen: announced via Hugging Face and company…
The developmentH Company announced the release of Holo4, a two-model series of open-weight computer-use agents, along with an updated Holotron4 Nano and open-sourced benchmark trajectories.

One Model Across Software Interfaces

Many workplace tasks combine actions that happen through different interfaces: a person might inspect a screen, write a short script and use a business API to finish one job. H Company argues that models built for only one interface can fail when a task crosses those boundaries. Holo4’s design aims to let one model choose between screen actions, code and tool calls.

If independent evaluations reproduce the reported results, the 27B model could offer developers a lower-cost option for some computer-use tasks, with the ability to run open weights in their own environments. That could matter to organizations that need to automate software but want more control over deployment and data handling. The release alone does not establish that Holo4 is reliable or economical on real business workloads; those outcomes depend on task performance, operating costs and the safeguards needed in practice.

H Company has also published the trajectories behind its public benchmark scores at trajectories.hcompany.ai and says they can be downloaded from Hugging Face. These records give outside evaluators material to inspect and reproduce. They make the claims more auditable, though they do not by themselves show that the benchmark setup reflects the variety of tasks businesses encounter.

Amazon

Top picks for "explor holo4 future"

As an affiliate, we earn on qualifying purchases.

From Holo Models to Holo4

Holo4 follows H Company’s earlier Holo1 agentic model and arrives alongside an updated model called Holotron4 Nano, according to the supplied announcement. The company’s benchmark notes identify Qwen3.8 27B as the dense model’s base and Qwen3.6 35B-A3B as the MoE model’s base.

H Company says it trained Holo4 to work with graphical interfaces as well as code and tools. It frames that combination as a response to the limits of single-interface agents: a model that depends on a visible screen may be unable to work when no screen is available, while a tool-calling model may have trouble with software that lacks an API. These are the company’s descriptions of the problem Holo4 is intended to address.

The announcement includes comparisons on OSWorld 2.0 and AutomationBench. H Company says its AutomationBench API-use result was measured in its internal harness, version 1.0.6. Its comparison draws on public-set scores for other models, while cost figures come from a leaderboard using the private set. The company says the models’ releases, harnesses and task subsets differ, limiting direct comparisons.

“Real work is not siloed that way, and a single business task can require combining these different approaches.”

— H Company

Benchmark Gaps and Open Questions

The headline scores remain self-reported by H Company. The source material does not provide independent reproductions of the OSWorld 2.0 results, nor does it establish how Holo4 performs across less curated business workflows. Although the company has released benchmark trajectories, outside evaluations are still needed to determine whether the published results hold under other harnesses and task selections.

The gap between the two Holo4 models is also unexplained: the 27B dense model is reported at 61.7% on OSWorld 2.0, compared with 30.9% for the larger 35B-A3B MoE model. The announcement material does not explain the difference. Holo4 has not yet been evaluated on AutomationBench’s private set, according to the source. The company’s cost comparisons also use different pricing and evaluation assumptions across models, so the figures do not establish a like-for-like cost advantage.

Independent Tests Will Follow

H Company says it plans to report Holo4’s results on the AutomationBench private set after that evaluation is complete. The next useful checks will include independent benchmark submissions and attempts to reproduce the OSWorld 2.0 results using the released weights and trajectories. Those evaluations can test whether the scores are repeatable and clarify how much results depend on a particular harness or task subset.

Developers can access Holo4 through the H Models API or download its weights from Hugging Face. Wider use may provide evidence about performance on everyday software tasks, but the announcement does not set out an adoption timeline or independent reliability findings. Until those checks are available, Holo4’s benchmark standing and practical cost advantage remain claims from its developer.

Key Questions

What is Holo4?

Holo4 is H Company’s series of open-weight agentic models, designed to operate software through graphical interfaces, code, MCP and APIs.

What models are included?

The release includes a 27B dense model and a 35B-A3B Mixture of Experts model. H Company says both are available through its API and for download in FP16, FP8 and GGUF formats.

How did Holo4 score on OSWorld 2.0?

H Company reports scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. The supplied material does not cite independent verification of these results.

Have the benchmark claims been independently verified?

The source material describes the scores as company-reported and does not cite independent reproductions. H Company has published the trajectories behind its public benchmark scores, which outside evaluators can inspect.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Your Intellectual Fly Is Open When You Use An LLM To Author A Post (2025)

Experts warn that employing large language models for content creation can unintentionally expose users’ lapses in judgment, akin to having an ‘open fly’.

The Impact Of GPT-5.6 Luna On AI Innovation And Software Development

OpenAI announces GPT-5.6 Luna integration with Replit to enhance AI-assisted software creation, but details on availability and capabilities remain unclear.

The LLM Critics Are Right. I Use LLMs Anyway

An analysis of why some users continue to rely on large language models despite ongoing criticism and concerns about their limitations.

15 AI Content Creation Platforms To Watch In 2026

Explore the top 15 AI content creation tools for 2026, highlighting features, strengths, and what makes each platform stand out for creators and marketers.