OpenAI & Anthropic
compatible inference

Point your existing SDK at this host, use your Lumid PAT as the API key, and call chat.completions or messages against a reasoning model — pick one with the model field (see /v1/models) or omit it for the default. Streaming on both shapes. Auth required — no PAT, no serving.

196K
context window
~25/s
output tok/s
~85/s
aggregate (8 streams)
6
endpoints

01Get a token

Every request authenticates with a Lumid Personal Access Token (PAT). It's the only credential you need — the same token is your API key for every SDK and curl call below.

1 · Sign inGo to lum.id and sign in (or create an account).
2 · Open TokensHead to lum.id/dashboard/tokens → Create token. Name it, set an expiry, and copy the value — it's shown once.
3 · Use itPass it as Authorization: Bearer <pat>, or as the api_key in any OpenAI / Anthropic SDK.
Treat a PAT like a password — it carries your identity and rate-limit tier. Store it in an env var, never commit it. Revoke and rotate from the same Tokens page if it leaks.

02Quick start

Model id is cyankiwi/MiniMax-M2.7-AWQ-4bit — or omit model and the server fills in the default.

Reasoning model — read this first. The model thinks before answering. The answer is in .reasoning (OpenAI shape) or the first thinking block (Anthropic shape) — .content is null / empty unless the model also emits a text reply. Set max_tokens ≥ 512 or the output is all reasoning with no answer.

OpenAI SDK

pip install openai
from openai import OpenAI

client = OpenAI(
    base_url="https://kv.run:5000/v1",
    api_key="<YOUR_LUMID_PAT>",
)
r = client.chat.completions.create(
    model="cyankiwi/MiniMax-M2.7-AWQ-4bit",
    messages=[{"role": "user",
               "content": "Hello"}],
    max_tokens=1024,
)
msg = r.choices[0].message
# answer is in reasoning; content may be null
print(msg.reasoning or msg.content)

Anthropic SDK

pip install anthropic
import anthropic

client = anthropic.Anthropic(
    base_url="https://kv.run:5000",
    api_key="<YOUR_LUMID_PAT>",
)
m = client.messages.create(
    model="cyankiwi/MiniMax-M2.7-AWQ-4bit",
    messages=[{"role": "user",
               "content": "Hello"}],
    max_tokens=1024,
)
# first block is thinking, last is the text reply
print(m.content[-1].text)

03Endpoints

All under /v1/. Add "stream": true for SSE on the generative endpoints.

GET /v1/modelsList all deployed model ids.
POST /v1/chat/completionsOpenAI chat completion. Streaming.
POST /v1/completionsOpenAI text completion. Streaming.
POST /v1/embeddingsOpenAI embeddings (if the model has an embedding head).
POST /v1/messagesAnthropic messages. Streaming.
POST /v1/messages/count_tokensAnthropic input-token count, no inference.
Available models. Default is cyankiwi/MiniMax-M2.7-AWQ-4bit (196K context). A 26B-parameter alternative is also available — pass its id from /v1/models in the model field to select it.

04Auth & limits

AuthRequired. Present a Lumid PAT as Authorization: Bearer <pat>. Anonymous requests get 401 — there is no free tier.
Rate limitPer-principal request cap, keyed on your PAT subject — limits are per-identity, not per-IP.
StreamingSSE pass-through. OpenAI stream=True and Anthropic messages.stream() both work unchanged.
Model fieldOptional — omit it and the default is injected. Pass cyankiwi/MiniMax-M2.7-AWQ-4bit to be explicit.
Errors504 upstream timeout · 502 backend unreachable · 503 no backend configured.