OpenAI-compatible inference with wallet auth and x402 credit billing. One endpoint, one key, every curated model.
No credit card required · free website chat
Get up and running with the OMA API in under a minute. Drop-in replacement for OpenAI with zero vendor lock-in.
Create an account and generate your API key from the dashboard.
Include your key in the Authorization header: Bearer oma_sk_...
Use the OpenAI SDK with base_url=https://www.oma-ai.com/api/v1
curl https://www.oma-ai.com/api/v1/chat/completions \
-H "Authorization: Bearer oma_sk_your_key_here" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.3-70b",
"messages": [{"role": "user", "content": "Hello!"}]
}'https://www.oma-ai.com/api/v1npm install openaiChat · Anthropic messages · curl
curl https://www.oma-ai.com/api/v1/chat/completions \
-H "Authorization: Bearer oma_sk_your_key_here" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.3-70b",
"messages": [{"role": "user", "content": "Hello!"}]
}'curl https://www.oma-ai.com/api/v1/messages \
-H "x-api-key: oma_sk_your_key_here" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"max_tokens": 256,
"messages": [{"role": "user", "content": "Hello!"}]
}'Programmatic inference uses a Bearer API key. Website accounts use email or wallet sessions; eligible free website chat also works without sign-in.
Eligible OMA-AI models can be used on /chat without an API key or wallet (20 msgs/day anonymous, 100/day signed-in), subject to live model availability. The chat model selector shows the current free options. Prepaid credits cover paid requests; rate limits still apply.
API-key calls (Authorization: Bearer oma_sk_…) always spend credits, including models eligible for free website chat. Free usage applies to browser website chat only, never to programmatic API traffic.
curl https://www.oma-ai.com/api/v1/chat/completions \
-H "Authorization: Bearer oma_sk_your_key_here" \
-H "Content-Type: application/json" \
-d '{"model":"llama-3.3-70b","messages":[{"role":"user","content":"Hello"}]}'Keys use the oma_sk_ prefix. Only the SHA-256 hash is stored. Keep your key secret and never commit it to version control.
Sign in with a passwordless magic link sent to your inbox. No password to leak, no password reset loops.
Connect any EVM wallet using Sign-In with Ethereum (SIWE). No email, no password. Your wallet address is your identity.
Use any OpenAI-compatible client: set base URL to www.oma-ai.com/api/v1 and Authorization: Bearer $OMA_KEY
One key, one base URL
https://www.oma-ai.com/api/v1oma_sk_…)llama-3.3-70b (see /models)POST /v1/messagesagent-skill.md file.Point OpenCode at OMA as a custom OpenAI-compatible provider. Env vars are the portable path; project config can pin the model.
# ~/.bashrc / shell profile (or OpenCode env UI)
export OPENAI_API_KEY="oma_sk_your_key_here"
export OPENAI_BASE_URL="https://www.oma-ai.com/api/v1"
# Optional: default model for sessions
export OPENAI_MODEL="llama-3.3-70b"
# Then run OpenCode as usual — it will hit OMA for chat completions.If your OpenCode build uses OPENAI_API_BASE instead of OPENAI_BASE_URL, set that to the same https://www.oma-ai.com/api/v1 value.
Project config (optional)
// opencode.json / provider settings (shape varies by OpenCode version)
{
"provider": {
"oma": {
"npm": "@ai-sdk/openai-compatible",
"name": "OMA",
"options": {
"baseURL": "https://www.oma-ai.com/api/v1",
"apiKey": "{env:OMA_API_KEY}"
},
"models": {
"llama-3.3-70b": { "name": "llama-3.3-70b" }
}
}
}
}Create the key at /dashboard → API keys. Never commit oma_sk_… to git — use env vars.
Claude Code talks the Anthropic Messages API. OMA exposes a compatible POST /v1/messages that rewrites to the same catalog. Configure a custom Anthropic base URL + key:
# Claude Code / Anthropic SDK env
export ANTHROPIC_API_KEY="oma_sk_your_key_here"
export ANTHROPIC_BASE_URL="https://www.oma-ai.com/api"
# Note: base URL is the origin (no /v1). The client appends /v1/messages.
# Model ids are OMA catalog ids (not anthropic/claude-3-5-sonnet):
# kimi-k3 · claude-opus-5 · llama-3.3-70b · …# One-shot smoke test (Messages API — x-api-key style)
curl https://www.oma-ai.com/api/v1/messages \
-H "x-api-key: oma_sk_your_key_here" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"max_tokens": 64,
"messages": [{"role": "user", "content": "ping"}]
}'
# Same call with Bearer (OpenAI-style auth also accepted)
curl https://www.oma-ai.com/api/v1/messages \
-H "Authorization: Bearer oma_sk_your_key_here" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{"model":"kimi-k3","max_tokens":64,"messages":[{"role":"user","content":"ping"}]}'Prefer one auth style consistently. Streaming: set stream: true — OMA returns Anthropic-shaped SSE. List models: GET /v1/models with the same key.
Any IDE that accepts an OpenAI-compatible endpoint:
https://www.oma-ai.com/api/v1oma_sk_…GET /v1/models// Example Continue config snippet (~/.continue/config.json)
{
"models": [
{
"title": "OMA llama-3.3-70b",
"provider": "openai",
"model": "llama-3.3-70b",
"apiBase": "https://www.oma-ai.com/api/v1",
"apiKey": "oma_sk_your_key_here"
}
]
}Drop-in replacement — only change base URL and key:
from openai import OpenAI
client = OpenAI(
api_key="oma_sk_your_key_here",
base_url="https://www.oma-ai.com/api/v1",
)
r = client.chat.completions.create(
model="llama-3.3-70b",
messages=[{"role": "user", "content": "Hello from OMA"}],
max_tokens=64,
)
print(r.choices[0].message.content)OpenAI-compatible REST API endpoints.
The public inference, model, balance, usage, and deposit endpoints.
OpenAI-compatible multi-turn chat. The endpoint you will hit 95% of the time. Supports SSE streaming.
{
"id": "chatcmpl-9f3k2a",
"object": "chat.completion",
"created": 1717194000,
"model": "llama-3.3-70b",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello. How can I help today?" },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 18, "completion_tokens": 12, "total_tokens": 30 }
}21 curated chat models across four tiers, plus web search in the playground. Check the live catalog before making a request.
Text and reasoning models from Venice (frontier/premium/mid/budget) and OMA-AI (MiniMax M2.7, DeepSeek V4 Flash, GLM-5.3 Flash).
A subset of premium / frontier models accept image inputs alongside text in the OpenAI multimodal content array shape.
Fast, cost-effective models for simple tasks and prototyping.
| Model | Input / 1M | Output / 1M | Context | Capabilities |
|---|---|---|---|---|
xiaomi-mimo-v2-5Xiaomi MiMo V2.5 | $0.50 | $2.50 | 1M | chatreasoningfast |
deepseek-v4-flash-0731DeepSeek V4 Flash (Venice) | $0.21875 | $0.4375 | 1M | chatreasoningfast |
mistral-small-2603Mistral Small 4 | $0.234375 | $0.9375 | 256K | chatfastagents |
google-gemma-4-31b-itGemma 4 31B | $0.15 | $0.45 | 256K | chatfast |
minimax-m2-7MiniMax M2.7 | $0.60 | $0.60 | 200K | chatreasoningcodingagents |
deepseek-v4-flash-0731-openbrokerDeepSeek V4 Flash (OMA-AI) | $0.50 | $0.50 | 400K | chatreasoningfastagents |
zai-org-glm-5-3GLM-5.3 Flash (OMA-AI) | $0.30 | $1.00 | 1M | chatreasoningmultimodalfastagents |
Balanced performance and cost for everyday workloads.
| Model | Input / 1M | Output / 1M | Context | Capabilities |
|---|---|---|---|---|
llama-3.3-70bLlama 3.3 70B | $0.875 | $3.50 | 128K | chatagents |
hermes-3-llama-3.1-405bHermes 3 405B | $1.375 | $3.75 | 128K | chatreasoningagents |
zai-org-glm-5-1GLM 5.1 | $1.925 | $6.05 | 200K | chatreasoningagents |
qwen-3-7-plusQwen 3.7 Plus | $0.625 | $2.50 | 1M | chatreasoningagents |
venice-uncensored-1-2Venice Uncensored 1.2 | $0.25 | $1.125 | 128K | chatuncensored |
Top-tier models for complex reasoning and generation tasks.
| Model | Input / 1M | Output / 1M | Context | Capabilities |
|---|---|---|---|---|
openai-gpt-56-solGPT-5.6 Sol | $7.8125 | $46.875 | 1M | chatreasoningcodingagents |
claude-opus-5Claude Opus 5 | $7.50 | $37.50 | 1M | chatreasoningcodingagents |
Use the GET /v1/models endpoint to fetch the current model catalog with pricing and capabilities.
curl https://www.oma-ai.com/api/v1/models | jq '.data | length'OMA routes each request through the best available provider — automatically.
Every model id resolves to an upstream provider. OMA routes each request to the model's catalog-mapped provider based on availability — with circuit-breaker protection against failing upstreams.
Primary provider for the curated chat catalog. Handles chat and multimodal-input workloads with low-latency inference.
Budget-tier models branded as OMA-AI — MiniMax M2.7, DeepSeek V4 Flash, and GLM-5.3 Flash on Gonka/OpenBroker. Same billing surface and API contract as the rest of the catalog.
Each response includes the public catalog provider label:
A mapped model may also have a compatible fallback provider configured. If a primary provider fails and that fallback is used, request content must be forwarded to the inference service that actually handles the retry. The public model id and listed customer price do not change. See Privacy for data-flow details.
Real-time token-by-token responses via Server-Sent Events (SSE).
Streaming provides real-time output as tokens are generated, reducing perceived latency from seconds to milliseconds. Essential for chat interfaces and interactive applications.
Add "stream": true to your request body. Responses arrive as Server-Sent Events (SSE) with Content-Type: text/event-stream.
curl https://www.oma-ai.com/api/v1/chat/completions \
-H "Authorization: Bearer oma_sk_your_key_here" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.3-70b",
"messages": [{"role": "user", "content": "Write a haiku about coding"}],
"stream": true
}'Each event contains a chat.completion.chunk with a delta field containing partial content. The stream ends with data: [DONE].
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1717194000,"model":"llama-3.3-70b","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1717194000,"model":"llama-3.3-70b","choices":[{"index":0,"delta":{"content":"Code"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1717194000,"model":"llama-3.3-70b","choices":[{"index":0,"delta":{"content":" flows"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1717194000,"model":"llama-3.3-70b","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]from openai import OpenAI
client = OpenAI(
api_key="oma_sk_your_key_here",
base_url="https://www.oma-ai.com/api/v1"
)
stream = client.chat.completions.create(
model="llama-3.3-70b",
messages=[{"role": "user", "content": "Write a haiku about coding"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")One set of limits for every key — no plan tiers. Enforced per API key (or per IP for unauthenticated requests).
| Limit | Value | Notes |
|---|---|---|
| Burst | 60 requests | Token bucket capacity — the most requests you can fire at once. |
| Refill | 20 req/s | How fast the bucket refills after a burst. |
| Sustained | 60 RPM | Sliding-window requests-per-minute ceiling. |
A few models carry their own request/token ceilings. Everything else uses the defaults above.
| Model | RPM | TPM |
|---|---|---|
llama-3.3-70b | 60 | 120,000 |
deepseek-v4-flash-0731-openbroker | 60 | 120,000 |
kimi-k3 | 60 | 120,000 |
llama-3.2-3b | 120 | 240,000 |
Token bucket algorithm allows short bursts above the sustained rate. The bucket holds 60 tokens and refills at 20 per second, so brief spikes succeed while sustained traffic settles at the RPM ceiling.
Per-model pricing. Input and output priced per million tokens.
| Type | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| GPT-5.6 Sol | $7.8125 | $46.875 | 1M |
| Claude Opus 5 | $7.50 | $37.50 | 1M |
| Kimi K3 | $4.6875 | $23.4375 | 1M |
| GLM 5.2 | $1.75 | $5.50 | 1M |
| DeepSeek V4 Pro | $2.0625 | $4.12625 | 1M |
| Qwen 3.8 Max | $3.125 | $9.375 | 1M |
| Grok 4.5 | $2.8375 | $8.50 | 500K |
| Claude Sonnet 4.6 | $4.50 | $22.50 | 1M |
| Gemini 3.6 Flash | $1.171875 | $5.859375 | 1M |
| Llama 3.3 70B | $0.875 | $3.50 | 128K |
| Hermes 3 405B | $1.375 | $3.75 | 128K |
| GLM 5.1 | $1.925 | $6.05 | 200K |
| Qwen 3.7 Plus | $0.625 | $2.50 | 1M |
| Venice Uncensored 1.2 | $0.25 | $1.125 | 128K |
| Xiaomi MiMo V2.5 | $0.50 | $2.50 | 1M |
| DeepSeek V4 Flash (Venice) | $0.21875 | $0.4375 | 1M |
| Mistral Small 4 | $0.234375 | $0.9375 | 256K |
| Gemma 4 31B | $0.15 | $0.45 | 256K |
| MiniMax M2.7 | $0.60 | $0.60 | 200K |
| DeepSeek V4 Flash (OMA-AI) | $0.50 | $0.50 | 400K |
| GLM-5.3 Flash (OMA-AI) | $0.30 | $1.00 | 1M |
Fund prepaid credits with a signed USDC authorization on Base.
Calling /x402/verify validates the EIP-3009 authorization, broadcasts the USDC transfer, and waits for confirmation. It is not a dry run. Sign in with the paying wallet before starting; /x402/settle credits only the SIWE session whose linked wallet matches the USDC payer.
Send base64-encoded JSON in PAYMENT-SIGNATURE. The v2 accepted fields specify the exact scheme, Base network, USDC asset, treasury, and atomic amount. The accepted amount and payTo must match authorization.value and authorization.to. The example below represents 1 USDC; replace every placeholder before signing.
{
"x402Version": 2,
"accepted": {
"scheme": "exact",
"network": "eip155:8453",
"asset": "0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913",
"amount": "1000000",
"payTo": "<configured treasury address>"
},
"payload": {
"signature": "<0x-prefixed 65-byte signature>",
"authorization": {
"from": "<payer wallet address>",
"to": "<same configured treasury address>",
"value": "1000000",
"validAfter": "<Unix seconds>",
"validBefore": "<Unix seconds>",
"nonce": "<0x-prefixed 32-byte random nonce>"
}
}
}Sign EIP-712 TransferWithAuthorization using the USDC domain: USD Coin, version 2, chain ID 8453, and the Base USDC contract. Times are Unix seconds, value is an integer string of atomic USDC units (6 decimals per token), and nonce is a fresh 32-byte value. Legacy raw-JSON PAYMENT-SIGNATURE envelopes and the older base64 X-PAYMENT header remain supported.
# Sign in via SIWE with the wallet that owns the USDC first.
# PAYMENT_SIGNATURE is a base64-encoded, signed v2 envelope.
# Calling verify broadcasts a transfer; it is not a dry run.
set -e
curl --fail-with-body -X POST https://www.oma-ai.com/x402/verify \
-H "Content-Type: application/json" \
-H "PAYMENT-SIGNATURE: $PAYMENT_SIGNATURE" > verification.json
jq -e '.valid == true' verification.json > /dev/null
# Credit the confirmed transfer to the same wallet's OMA account.
# Reuse the returned txHash and networkId; do not submit an amount.
jq '{txHash, networkId}' verification.json | \
curl --fail-with-body https://www.oma-ai.com/x402/settle \
-H "Content-Type: application/json" \
--cookie "$OMA_SESSION_COOKIE" \
--data-binary @-
# A successful response includes settled, receipt, and creditedCents.
# Now use that account's API key to spend the prepaid balance.These shell examples require curl and jq. Keep private keys, session cookies, and signed payment envelopes out of logs and source control. The public funding URLs use /x402/*; the SDK base URL is https://www.oma-ai.com/api/v1.
If funds moved but credits are pending, retry /x402/settle with the same hash and payer session. Crediting is idempotent and can recover the payer and amount from the confirmed USDC Transfer log after a server restart. A successful response includes a signed receipt string and creditedCents; PAYMENT-RESPONSE carries a base64 receipt summary.
A verification timeout can include a txHash even when valid is false. Check its Base receipt before authorizing another payment. Once confirmed, settle that hash. If no hash was returned, reconcile the signed nonce and wallet activity first.
Base mainnet only (eip155:8453), using USDC. The service pins the USDC contract and configured treasury. Settlement accepts networkId base or eip155:8453 and derives the amount from verified payment evidence rather than a caller-supplied amount.
Complete, working examples in every major language. OMA is fully OpenAI-compatible — just swap the base URL.
pip install openainpm install openaigo get github.com/sashabaranov/go-openaiBuilt-in (no install)Send a multi-turn chat request with system prompt, temperature control, and token limits.
from openai import OpenAI
client = OpenAI(
api_key="oma_sk_your_key_here",
base_url="https://www.oma-ai.com/api/v1"
)
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing in one sentence."}
],
temperature=0.7,
max_tokens=256
)
print(response.choices[0].message.content)Real-time token-by-token responses via Server-Sent Events. Add "stream": true to your request.
from openai import OpenAI
client = OpenAI(
api_key="oma_sk_your_key_here",
base_url="https://www.oma-ai.com/api/v1"
)
stream = client.chat.completions.create(
model="llama-3.3-70b",
messages=[{"role": "user", "content": "Write a haiku about coding"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")Fetch the full model catalog with pricing and capabilities. No auth required for the public endpoint.
from openai import OpenAI
client = OpenAI(
api_key="oma_sk_your_key_here",
base_url="https://www.oma-ai.com/api/v1"
)
models = client.models.list()
for model in models.data:
print(f"{model.id} ({model.get('category', 'chat')})")All examples use the same base URL. Just swap your API key and you're ready:
https://www.oma-ai.com/api/v1For Next.js and React apps, use the Vercel AI SDK with OMA as the provider:
import { generateText } from 'ai';
import { createOpenAI } from '@ai-sdk/openai';
const oma = createOpenAI({
baseURL: 'https://www.oma-ai.com/api/v1',
apiKey: 'oma_sk_your_key_here',
});
const { text } = await generateText({
model: oma('llama-3.3-70b'),
prompt: 'Explain quantum computing in one sentence.',
});Standard HTTP status codes returned by the API.
Every error returns a JSON body with an error object carrying message, type, and code.
{
"error": {
"message": "Rate limit exceeded. Please slow down.",
"type": "rate_limit_error",
"code": "rate_limit_exceeded"
}
}| Code | Status | Description |
|---|---|---|
| 400 | invalid_request_error | Malformed request body or invalid parameters. Schema failures return validation_error. |
| 401 | authentication_error | Missing or invalid API key. |
| 402 | payment_error | Insufficient credits. Deposit more USDC or settle via x402, then retry. |
| 403 | forbidden | Authenticated but not allowed. E.g. a resource/endpoint you lack permission for, or feature_disabled. |
| 404 | not_found | Route, model, key, or resource does not exist. |
| 409 | conflict | Request conflicts with current state (e.g. settle already applied to another account). |
| 422 | validation_error | Request is well-formed but semantically invalid; also returned as invalid_request_error for schema failures. |
| 429 | rate_limit_error | Rate limit exceeded. Wait Retry-After seconds, then retry. Specific codes: rate_limit_exceeded, daily_limit_exceeded. |
| 500 | internal_error | Unexpected server error (also server_error). Contact support if persistent. |
| 503 | service_unavailable | All upstream providers down. Try again shortly. |
import time
from openai import OpenAI
client = OpenAI(api_key="oma_sk_...", base_url="https://www.oma-ai.com/api/v1")
for attempt in range(5):
try:
response = client.chat.completions.create(model="llama-3.3-70b", messages=[{"role": "user", "content": "Hi"}])
break
except Exception as e:
if hasattr(e, 'status_code') and e.status_code == 429:
time.sleep(2 ** attempt)
else:
raise