Skip to main content
GET
List models
Lists the models currently served by the API, with capabilities, limits, pricing, and lifecycle metadata. Use it to discover model IDs at runtime and to detect upcoming removals via deprecation_date instead of hardcoding assumptions. Returns a list object whose data array contains one object per model.

Response

string
The model ID to pass as model in chat completions, e.g. holo4-27b.
integer
Total context window in tokens.
integer
Hard ceiling on output tokens per request. 8,192 on the Holo3 models.
array
["text", "image"] for all Holo models.
array
Capability flags: reasoning on every model, and tools for native function calling on every model except holo3-122b-a10b.
array
Accepted sampling fields, e.g. temperature, top_p, top_k, max_tokens, stop, frequency_penalty, presence_penalty, seed.
object
Per-token USD rates as decimal strings: prompt and completion per input/output token, and input_cache_read per prompt token served from the cache. See Prompt caching.
string
Set when the model is scheduled for removal; null otherwise. After removal, requests to the ID fail with model_not_found.
boolean
Whether the model is currently serving traffic (also see is_ready).

Examples

Response (truncated)

Next steps

Chat completions

The endpoint these model IDs are used with.

Models API

Model cards and pricing.