H Company has released Holo4, a family of generalist computer-use models for AI agents. One set of weights clicks and types on screens. It also writes code and calls MCP or API tools. Holo4 ships in 2 sizes: Holo4 27B (dense) and Holo4 35B-A3B (Mixture of Experts, 3B active). Both serve a 256K context on the H Models API.
Is it deployable? Yes. Holo4 35B-A3B ships Apache 2.0 weights for commercial self-hosting. Holo4 27B weights are CC BY-NC 4.0, so commercial use of 27B runs through the H Models API.
What is Holo4
Holo4 is a vision-language model for computer use. Holo4 27B is fine-tuned from Qwen3.8-27B. Holo4 35B-A3B is built on Qwen3.6-35B-A3B. Both pair with H’s open hai-agents harness. The harness sends screenshots and tool results to the model. It then executes the requested clicks, typing, code and tool calls.
H company targets a known gap. GUI-only agents fail without a screen. Tool-calling agents stall when an application has no API. Holo4 runs on desktop, web, Android, code sandboxes and business APIs. It is the same model, called the same way, on every platform.
Benchmarks: Close to the Frontier, at a Fraction of the Cost
Per H Company’s benchmark table, Holo4 27B scores 85.2% on OSWorld at $0.08 per task. Its Qwen3.8 27B base scores 84.3% at $0.22. On AndroidWorld, Holo4 27B reaches 85.1%.
Long workflows show the remaining gap. On OSWorld 2.0, Holo4 27B scores 61.7% at $1.22 per task. Claude Opus 5.5 scores 81.8% at $8.48, per H’s figures. On AutomationBench, Holo4 27B scores 45.4% at $0.05 per task.
It is important to note that frontier scores come from different harnesses and effort levels. Also, 480 of AutomationBench’s 600 public tasks sit in the split H Company collected training data from. On the 120 held-out tasks, Holo4 27B scores 49.3%. H Company publishes every trajectory at trajectories.hcompany.ai and on Hugging Face.
How H Company Built Holo4
Agentic Task Factory: H’s internal pipelines build environments and verifiable tasks from documentation, screenshots and real software. The factory has produced about 10,000 tasks: 4k web apps, 3k MCP servers, 3k desktop and OS. A task survives only if its verifier rejects near misses. An agent must also solve it through the real interface.
Supervised fine-tuning: The SFT set holds 127B tokens. About three quarters are successful agentic trajectories: desktop 45%, web 14%, MCP and API 12%, mobile 3%.
2 RL experts, 1 merge: Asynchronous online RL trains 2 LoRA experts. One handles desktop and web. The other handles terminal, MCP and API. Both merge back with equal weight and no further training.
Harness: H Company rebuilt its agent loop using OSWorld 2.0 failure analysis. The largest changes were reliable memory across hundreds of steps and a shell on the desktop machine.
Holotron4 Nano
H Company also released Holotron4 Nano, built on NVIDIA’s Nemotron 3 Nano Omni through the Nemotron Coalition. The same combination lifts OSWorld from 21.0% to 76.3% over the base model.
Pricing and Deployment
Holo4 27B costs $0.40 input and $3.00 output per 1M tokens. Holo4 35B-A3B costs $0.30 and $2.00. The API is OpenAI-compatible at https://api.hcompany.ai/v1; see the quickstart. Weights on the Hugging Face collection come in BF16, FP8, NVFP4 and 4-bit GGUF. H Company documents local inference with vLLM and llama.cpp. H says DSpark drafter checkpoints for faster inference arrive in the coming days.
Interactive Explainer: How Holo4 Works
‘;resize()}
function line(h){var d=document.createElement(‘div’);d.innerHTML=h;$(‘#log’).appendChild(d);var L=$(‘#log’);while(L.children.length>9)L.removeChild(L.firstChild)}
function step(){var s=tasks[ti].seq;if(si>=s.length){clearInterval(timer);timer=null;return}
var a=s[si];clearAct();
var phase=si===0?0:1;$$(‘.st’)[0].classList.add(‘on’);
if(a[0]===’done’){$$(‘.st’)[3].classList.add(‘on’);line(‘done ‘+a[2])}
else{$(‘#n-‘+a[0]).classList.add(‘act’);$(‘#w-‘+a[0]).classList.add(‘act’);$$(‘.st’)[1].classList.add(‘on’);$$(‘.st’)[2].classList.add(‘on’);
var lab={gui:’GUI’,code:’CODE’,tool:’MCP’,api:’API’}[a[0]];
line(‘observe screenshot + last result’);
line(‘‘+lab+’.’+a[1]+’ ‘+a[2])}
si++;resize()}
$(‘#run’).onclick=function(){reset();$(‘#log’).innerHTML=”;step();timer=setInterval(function(){step();if(si>=tasks[ti].seq.length){clearInterval(timer);timer=null}},1300)};
$(‘#nxt’).onclick=function(){if(si===0)$(‘#log’).innerHTML=”;step()};
reset();
/* PANE 2 */
var mix=[[‘Desktop’,45,’#111111′,’#F0EEE9′],[‘Web’,14,’#4F3BFF’,’#fff’],[‘MCP and API’,12,’#D8F25C’,’#111′],[‘Mobile’,3,’#8C7DFF’,’#fff’],[‘Other’,26,’#D6D2CA’,’#111′]];
mix.forEach(function(m){var s=document.createElement(‘div’);s.className=”seg”;s.style.background=m[2];s.style.color=m[3];s.dataset.w=m[1];s.textContent=m[1]>=10?m[1]+’%’:”;$(‘#mix’).appendChild(s);
var l=document.createElement(‘span’);l.innerHTML=’‘+m[0]+’ ‘+m[1]+’%’;$(‘#leg’).appendChild(l)});
function drawMix(){$$(‘.seg’).forEach(function(s){s.style.width=”0″});setTimeout(function(){$$(‘.seg’).forEach(function(s){s.style.width=s.dataset.w+’%’})},60)}
var stages=[
[’01’,’Agentic Task Factory’,’About 10,000 verifiable tasks built from docs, screenshots and real software: 4k web apps, 3k MCP servers, 3k desktop and OS.
- Verifier must fail on the untouched seed
- Pass on the golden state and reject every near miss
- An agent must solve it through the real interface
‘],
[’02’,’Supervised fine-tuning’,’127B tokens. About three quarters are successful agentic trajectories from the Task Factory. The rest covers multimodal reasoning, GUI grounding, and text-only tool use and coding.’],
[’03’,’2 RL experts’,’Asynchronous online RL on long-horizon tasks trains 2 LoRA experts on the fine-tuned model:
- Expert A: desktop and web
- Expert B: terminal, MCP and API
‘],
[’04’,’Merge’,’Both experts merge back into the fine-tuned model with equal weight and no further training. Output: one Holo4 checkpoint for every interface.’]];
stages.forEach(function(s,i){var d=document.createElement(‘div’);d.className=”stage”+(i?”:’ on’);d.innerHTML=’
‘+s[0]+’
‘+s[1]+’
‘;d.onclick=function(){$$(‘.stage’).forEach(function(x){x.classList.remove(‘on’)});d.classList.add(‘on’);$(‘#det’).innerHTML=s[2];resize()};$(‘#pipe’).appendChild(d)});
$(‘#det’).innerHTML=stages[0][2];
/* PANE 3 */
var B={
‘OSWorld’:{rows:[[‘Holo4 27B’,85.2,0.08,’h’],[‘Holo4 35B-A3B’,80.8,0.05,’h’],[‘Qwen3.8 27B (base)’,84.3,0.22,’q’],[‘Fable 5′,86.0,null,’f’],[‘GPT-5.5′,78.7,null,’f’]],note:’Computer tasks. Holo4 in H harness, mean of 2 to 4 runs. Fable 5 from the official leaderboard; GPT-5.5 as reported by OpenAI.’},
‘OSWorld 2.0’:{rows:[[‘Holo4 27B’,61.7,1.22,’h’],[‘Holo4 35B-A3B’,30.9,0.61,’h’],[‘Qwen3.8 27B (base)’,48.0,3.49,’q’],[‘Claude Opus 5.5′,81.8,8.48,’f’],[‘GPT-6 Astra’,73.5,9.07,’f’]],note:’Long computer workflows, average partial score, single Holo4 run. Opus 5.5 at max effort in Anthropic harness; GPT-6 Astra at max effort on the 82-task offline subset.’},
‘AutomationBench’:{rows:[[‘Holo4 27B’,45.4,0.05,’h’],[‘Holo4 35B-A3B’,34.5,0.02,’h’],[‘Qwen3.8 27B (base)’,40.3,0.09,’q’],[‘Claude Opus 5′,50.3,3.05,’f’],[‘Kimi K3′,46.7,0.43,’f’],[‘GPT-5.6 Sol’,45.8,0.67,’f’]],note:’Business automation, 600 public v1.0.6 tasks. 480 fall in the split H collected training data from. On the 120 held-out tasks: Holo4 27B 49.3%, 35B-A3B 31.7%.’},
‘AndroidWorld’:{rows:[[‘Holo4 27B’,85.1,0.08,’h’],[‘Holo4 35B-A3B’,77.6,0.07,’h’],[‘Qwen3.8 27B (base)’,81.9,0.13,’q’],[‘Fable 5′,88.8,null,’f’],[‘Qwen3.8 Max’,85.3,null,’f’],[‘GPT-5.6 Sol’,77.6,null,’f’]],note:’Phone apps. Frontier scores as measured by Qwen for the Qwen3.8 release.’},
‘ALE-CLI’:{rows:[[‘Holo4 27B’,44.1,0.82,’h’],[‘Holo4 35B-A3B’,30.9,0.29,’h’],[‘Qwen3.8 27B (base)’,43.5,null,’q’],[‘Claude Opus 5.5′,63.7,8.22,’f’],[‘GPT-6 Astra’,61.4,5.31,’f’],[‘Muse Spark 1.3′,57.5,2.44,’f’]],note:’Expert Linux workflows (105-task split of Agents\u2019 Last Exam), average score. Frontier numbers from the official leaderboard.’}
};
var col={h:’#4F3BFF’,q:’#8C7DFF’,f:’#111111′},cur=”OSWorld”;
Object.keys(B).forEach(function(k,i){var c=document.createElement(‘span’);c.className=”chip”+(i?”:’ on’);c.textContent=k;c.onclick=function(){$$(‘#bench .chip’).forEach(function(x){x.classList.remove(‘on’)});c.classList.add(‘on’);cur=k;drawBench(k)};$(‘#bench’).appendChild(c)});
function drawBench(k){var b=B[k],h=””;
b.rows.forEach(function(r){h+=’
‘+r[0]+’
‘+r[1].toFixed(1)+’%‘+(r[2]!==null?’$’+r[2].toFixed(2)+’/task’:’cost n/a’)+’
‘});
$(‘#bars’).innerHTML=h;setTimeout(function(){$$(‘.fl’).forEach(function(f){f.style.width=f.dataset.w+’%’})},50);
var hol=b.rows[0],top=b.rows.slice().sort(function(a,c){return c[1]-a[1]})[0],k1,k2;
k1=’
‘+(hol[1]/top[1]*100).toFixed(0)+’%of the top score (‘+top[0]+’) reached by Holo4 27B
‘;
var ref=b.rows.filter(function(r){return r[3]===’f’&&r[2]!==null}).sort(function(a,c){return c[2]-a[2]})[0];
k2=ref?’
‘+(ref[2]/hol[2]).toFixed(1)+’xhigher cost per task for ‘+ref[0]+’ vs Holo4 27B
‘:’
$’+hol[2].toFixed(2)+’Holo4 27B cost per task at H Models API rates
‘;
$(‘#kpi’).innerHTML=k1+k2;
$(‘#bnote’).innerHTML=b.note+’ Public frontier scores use different harnesses and effort levels, per H Company. Costs: H Models API rates for Holo4, Alibaba Cloud list prices for Qwen3.8 27B, provider charts for frontier models.’;
resize()}
drawBench(cur);
/* PANE 4 */
var A={};
$$(‘.opt’).forEach(function(o){o.onclick=function(){$$(‘.opt[data-q=”‘+o.dataset.q+'”]’).forEach(function(x){x.classList.remove(‘on’)});o.classList.add(‘on’);A[o.dataset.q]=o.dataset.v;rec()}});
function rec(){if(!A.host||!A.com||!A.pri){$(‘#rec’).innerHTML=”;resize();return}
var t,d,c;
if(A.host===’api’){
if(A.pri===’acc’){t=”Holo4 27B on the H Models API”;d=’Flagship dense model, 256K context. $0.40 input / $3.00 output per 1M tokens ($0.04 cached). Commercial use is allowed through the API.’;c=”model=”holo4-27b””}
else{t=”Holo4 35B-A3B on the H Models API”;d=’H\u2019s default for production computer use. 3B active parameters, 256K context. $0.30 input / $2.00 output per 1M tokens ($0.03 cached).’;c=”model=”holo4-35b-a3b””}
c=”from openai import OpenAI\nclient = OpenAI(base_url=”https://api.hcompany.ai/v1″, api_key=HAI_API_KEY)\nclient.chat.completions.create(“+c+’, messages=[…])’}
else if(A.com===’y’){t=”Self-host Holo4 35B-A3B (Apache 2.0)”;d=(A.pri===’acc’?’Holo4 27B weights are CC BY-NC 4.0, so they cannot power a commercial product off-API. ‘:”)+’The 35B-A3B weights are Apache 2.0. Pick BF16, FP8, NVFP4 or 4-bit GGUF and serve with vLLM or llama.cpp.’;c=”vllm serve Hcompany/Holo4-35B-A3B-FP8 \\\n –served-model-name holo4-35b-a3b”}
else{if(A.pri===’acc’){t=”Self-host Holo4 27B (CC BY-NC 4.0)”;d=’Non-commercial use of the 27B weights is permitted. Available in BF16, FP8, NVFP4 and Q4 GGUF.’;c=”vllm serve “Hcompany/Holo4-27B””}
else{t=”Self-host Holo4 35B-A3B”;d=’Only 3B parameters active per token. Apache 2.0, with BF16, FP8, NVFP4 and GGUF builds.’;c=”vllm serve Hcompany/Holo4-35B-A3B-FP8 \\\n –served-model-name holo4-35b-a3b”}}
$(‘#rec’).innerHTML=”;resize()}
window.addEventListener(‘load’,resize);window.addEventListener(‘resize’,resize);setTimeout(resize,300);
})();
