Point your existing SDK at this host, use your Lumid PAT as the API key,
and call chat.completions or messages against a
reasoning model — pick one with the model field
(see /v1/models) or omit it for the default. Streaming on both shapes.
Auth required — no PAT, no serving.
Every request authenticates with a Lumid Personal Access Token (PAT). It's the only credential you need — the same token is your API key for every SDK and curl call below.
| 1 · Sign in | Go to lum.id and sign in (or create an account). |
| 2 · Open Tokens | Head to lum.id/dashboard/tokens → Create token. Name it, set an expiry, and copy the value — it's shown once. |
| 3 · Use it | Pass it as Authorization: Bearer <pat>, or as the api_key in any OpenAI / Anthropic SDK. |
Model id is cyankiwi/MiniMax-M2.7-AWQ-4bit — or omit model and the
server fills in the default.
.reasoning (OpenAI shape) or the first
thinking block (Anthropic shape) — .content is
null / empty unless the model also emits a text reply.
Set max_tokens ≥ 512 or the output is all reasoning with no answer.
from openai import OpenAI client = OpenAI( base_url="https://kv.run:5000/v1", api_key="<YOUR_LUMID_PAT>", ) r = client.chat.completions.create( model="cyankiwi/MiniMax-M2.7-AWQ-4bit", messages=[{"role": "user", "content": "Hello"}], max_tokens=1024, ) msg = r.choices[0].message # answer is in reasoning; content may be null print(msg.reasoning or msg.content)
import anthropic client = anthropic.Anthropic( base_url="https://kv.run:5000", api_key="<YOUR_LUMID_PAT>", ) m = client.messages.create( model="cyankiwi/MiniMax-M2.7-AWQ-4bit", messages=[{"role": "user", "content": "Hello"}], max_tokens=1024, ) # first block is thinking, last is the text reply print(m.content[-1].text)
All under /v1/. Add "stream": true for SSE
on the generative endpoints.
| GET /v1/models | List all deployed model ids. |
| POST /v1/chat/completions | OpenAI chat completion. Streaming. |
| POST /v1/completions | OpenAI text completion. Streaming. |
| POST /v1/embeddings | OpenAI embeddings (if the model has an embedding head). |
| POST /v1/messages | Anthropic messages. Streaming. |
| POST /v1/messages/count_tokens | Anthropic input-token count, no inference. |
cyankiwi/MiniMax-M2.7-AWQ-4bit (196K context). A 26B-parameter
alternative is also available — pass its id from /v1/models in the
model field to select it.
| Auth | Required. Present a Lumid PAT as Authorization: Bearer <pat>. Anonymous requests get 401 — there is no free tier. |
| Rate limit | Per-principal request cap, keyed on your PAT subject — limits are per-identity, not per-IP. |
| Streaming | SSE pass-through. OpenAI stream=True and Anthropic messages.stream() both work unchanged. |
| Model field | Optional — omit it and the default is injected. Pass cyankiwi/MiniMax-M2.7-AWQ-4bit to be explicit. |
| Errors | 504 upstream timeout · 502 backend unreachable · 503 no backend configured. |