Holo3.1: Fast & Local Computer Use Agents
HCompany says Holo3.1 broadens its computer-use model to desktop, web, and mobile, while adding local quantized checkpoints.
Intelligence analysis by GPT-5.4 Mini

Holo3.1 is pitched as a production-ready update to HCompany’s computer-use agents. The release focuses on three things: better cross-environment robustness, support for different agent harnesses, and local deployment through quantized checkpoints.
Holo3.1 is like giving a robot helper better shoes for different floors, plus a version that can live inside a laptop instead of always calling home. HCompany says that makes the helper work better on phones, desktops, and private computers.
Analysis
What HCompany is shipping
HCompany says Holo3.1 is the next version of its computer-use model family, built on Qwen and aimed at making agents more reliable across web, desktop, and mobile environments. The company frames the release as a response to a practical issue it saw after Holo3 adoption grew: strong results in one setup did not always transfer cleanly to another.
Main improvements
The article highlights three production concerns. First, Holo3.1 improves mobile automation, with the 35B-A3B model moving from 67% to 79.3% on AndroidWorld, while the 4B and 9B models rise from 58% to 72%. Second, it adds native function-calling support alongside the structured JSON outputs already available in Holo3, and HCompany says this brings near-parity performance across OSWorld and its internal benchmark suite. Third, the company says Holo3.1 improves performance inside its Holotab harness by more than 25% versus Holo3.
Local and private deployment
HCompany is also pushing smaller model sizes for cost and privacy tradeoffs: 0.8B, 4B, 9B, and 35B-A3B. This is the first Holo release with quantized checkpoints, including FP8, Q4 GGUF, and NVFP4 for the 35B-A3B model. The company says these checkpoints are intended to enable fast local inference with little to no loss in model quality, and that FP8 and NVFP4 land about two points below BF16 on OSWorld.
The article also says NVFP4 W4A16 on DGX Spark reaches 1.41x the token throughput of FP8 and 1.74x BF16, and that on Spark the combined harness and quantization work cuts average step time from 6.8 seconds to 3.3 seconds. HCompany says the agent can run locally on Windows or Mac, with the model either on the same machine or on a nearby DGX Spark, keeping execution private and inside the user’s network.
Bottom line
Holo3.1 is positioned less as a flashy benchmark release and more as an engineering package for practical deployment: better across environments, more compatible with agent stacks, and more usable on local hardware.
Key points
- Holo3.1 is presented as a computer-use model family built for web, desktop, and mobile workflows.
- HCompany says mobile performance improved sharply on AndroidWorld across the 35B-A3B, 4B, and 9B models.
- The release adds native function-calling support and claims near-parity with structured JSON outputs in key benchmarks.
- This is the first Holo release with quantized checkpoints, including FP8, Q4 GGUF, and NVFP4.
- HCompany says the local setup is designed for private execution on end-user devices or nearby hardware.
If HCompany’s numbers hold up in real use, teams could get computer-use agents that work across web, desktop, and mobile without rewriting their stacks. The local quantized checkpoints could also make private deployment more practical for organizations that do not want every action sent to the cloud.
The article’s gains are benchmark-driven, so real-world workflows may still behave differently once teams move beyond the tested environments. Local and mobile deployments can also be limited by hardware constraints, especially for larger models or if the deployment stack does not match the company’s reference setup.



