Quickstart for Minicpm users
Connect your existing OpenAI-compatible client to our uncensored API in under a minute. Replace your base URL, set your key, and start generating text with the "uncensored" model.
Base URL and Authentication
To use this API, you need an API key and the correct base URL. Our service is fully compatible with the standard OpenAI client libraries. You do not need to install new dependencies or learn a new SDK. Simply point your existing client to our endpoint.
Generate a key on the Get API key page. You can sign in with Google or use an email and password. The key is displayed immediately. Store it securely. Every request must include this key in the Authorization header.
The base URL for all requests is: https://api.minicpm.cc/v1. This URL handles chat completions, model listing, and streaming responses. There are no separate endpoints for embeddings or vision. This is a text-only API.
First Request
Send a standard chat completion request to start generating text. The model ID is always uncensored. This model is an open-weight large language model tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or any other vendor's model.
Set the model field to uncensored. Provide your messages in the standard array format. The API returns a text response. You can control the output length with max_tokens. If you do not set it, the default max output is 2,048 tokens.
Python SDK Integration
Use the official OpenAI Python library to interact with the API. You only need to configure the base URL and API key. The rest of the library works as expected. This includes support for streaming, function calling, and JSON mode.
Install the library with pip install openai. Initialize the client with your base URL and key. Send messages to the chat completions endpoint. The response object contains the generated text, token usage, and finish reason. You can access these fields directly in your Python code.
Node SDK Integration
The Node.js SDK works identically to the Python version. Install the package and configure the base URL. The client handles authentication and request formatting automatically. This allows you to integrate the API into your backend services quickly.
Pass the base URL and API key to the OpenAI client constructor. Use the chat.completions.create method. The response structure matches the OpenAI standard. You can handle errors, check token usage, and process the content. This approach ensures compatibility with existing codebases that already use OpenAI.
Streaming Responses
Enable streaming by setting stream: true in your request. The API returns a sequence of Server-Sent Events (SSE). Each event contains a chunk of the response. This allows you to display text to the user as it is generated.
The final chunk includes the token usage statistics. This is useful for tracking costs. Streaming reduces the perceived latency for users. It is recommended for chat interfaces. You can parse the events in your client code to build the final response or display tokens in real time.
Limits, Errors, and Context
The model supports a 100,000 token context window. This includes both the prompt and the completion. The maximum output per request is 32,000 tokens. You can adjust the max_tokens parameter to control this.
Rate limits are set to 300 requests per minute per key. You can have 8 concurrent requests. The request body must be under 8 MB. If the key is invalid, you will receive a 401 error. If you have no credit, you will receive a 402 error. If you exceed the rate limit, you will receive a 429 error. Errors and refusals do not consume credit.
cURL
curl https://api.minicpm.cc/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'Python
from openai import OpenAI
client = OpenAI(base_url="https://api.minicpm.cc/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Node.js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.minicpm.cc/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Streaming
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Under the hood: specs
If your tool speaks the OpenAI API, these are the details that matter.
| Spec | Value |
|---|---|
| Compatibility | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Base URL | https://api.minicpm.cc/v1 |
| Model ID | uncensored |
| Authentication | Authorization: Bearer YOUR_KEY |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Function calling | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| JSON mode | JSON object mode via response_format json_object |
| Max context | 100,000 tokens, input and output combined |
| Completion length | up to the rest of the 100,000-token window; max_tokens optional (no separate cap) |
| Concurrency | up to 8 in parallel per key |
| Max body | 8 MB request body |
| Rate limit | 300/min per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| How you pay | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Payment | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Subscription | no monthly fee; paid credit does not expire |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Free trial | $0.50 for 7 days, no card · Trial key: 2 parallel requests, 60 req/min; full limits (8 and 300) after first top-up |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Sign-in | sign in with Google or with e-mail + password |
| Content | adult content allowed; sexual content involving minors is refused |
| Keys | one active key per account; a new key replaces the old one |
Error codes
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | unknown endpoint |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What is the context window size?
The context window is 100,000 tokens, which includes both the input prompt and the output completion. The maximum output per request is 32,000 tokens.
How do I handle errors?
A 401 error indicates an invalid API key. A 402 error indicates that your prepaid credit is exhausted. A 429 error means you have exceeded the rate limit of 300 requests per minute. Errors do not consume your credit.
Can I use streaming?
Yes. Set <code>stream: true</code> in your request. The API returns Server-Sent Events (SSE). The final chunk contains the token usage statistics.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.