Skip to main content
The H Company Models API gives developers access to the Holo Vision-Language Models: multimodal models trained to operate real user interfaces across web, desktop, and mobile. Through a single OpenAI-compatible API, you send text, images, or both, and receive structured outputs you can act on directly. Two models are served today: the fast, open-weight holo3-1-35b-a3b (free tier) and the maximum-performance holo3-122b-a10b. See Models for capabilities, limits, and pricing.

Two ways to use Holo

There is also a third, non-GUI pattern: Document OCR, the same endpoint used as a one-shot page transcriber.

Get started

Quickstart

First request in five minutes.

Models

What is served, limits, and pricing.

API reference

Endpoint, conventions, and parameters.

Agent loop

Use Holo in your computer-use harness.

Model cards and benchmarks

Hugging Face

Model cards, weights, and quantized builds.

Holo3.1 release

Mobile, function calling, and local inference.

Holo3 benchmarks

78.85% on OSWorld-Verified.
Prefer to try the models without writing code? HoloTab runs Holo directly in your browser without any setup.