Skip to main content
NewHolo4 is here. One model for screens, code, and tools.Read the announcement → Holo is a family of agent models. Hand Holo4 a toolbox and it works through whatever the task needs: it clicks and types on a screen, writes and runs code, or calls your APIs and MCP servers, often in the same run. The API is OpenAI-compatible, and every interface uses the same call: tools in, tool_calls out.

Models

Start with holo4-35b-a3b. It has a 262,144-token context, and its latency suits interactive loops. On the free tier, use holo3-1-35b-a3b. Move to holo4-27b for complex multi-step tasks and novel environments. The Holo3 generation stays available. Switching is a model ID change. IDs are stable: before a model is removed, its deprecation_date is set in GET /v1/models and a notice appears on this page, so pin an ID in production and check deprecation_date when you upgrade.

What Holo does

Prefer to try Holo without writing code? HoloTab runs it in your browser with no setup.

Benchmarks

Scores, and how each is measured, are in the release posts: Holo4 and Holo3.1.

Smaller sizes for your own hardware

Both Holo4 models are open weights, in BF16, FP8, NVFP4, and Q4 GGUF; each model card on Hugging Face states its license. Three smaller Holo3.1 checkpoints are not on the API, but they are open weights under Apache 2.0: download them and run them yourself. Holo3.1 35B is open too, with FP8, NVFP4, and GGUF builds for a Mac (M3 or newer, 36 GB) or a DGX Spark. Local inference covers memory, measured speed, and server setup with llama.cpp or vLLM.

Data retention

The API logs request time, model, and token counts. Prompts and responses are not stored: zero data retention by default. H runs its own models, so inputs and outputs are never shared with third parties. See the privacy policy.