Quickstart: Integrate the Uncensored LLM API
Integrate our uncensored AI chatbot API into your application using standard OpenAI-compatible endpoints. This quickstart covers authentication, basic requests, streaming, and function calling with our single uncensored model.
Authentication and Base URL
Our API uses standard OpenAI-compatible authentication. You need a valid API key, which you can generate immediately on the Get API key page. The base URL for all requests ishttps://api.unrestrictedaichatbot.cc/v1. Configure your client by setting this base URL and your API key. We support Google login or email/password registration. No phone number is required. Your API key is shown once after creation. Keep it secure. Each account allows one active key; generating a new one replaces the old. Requests are authenticated via the Authorization: Bearer <API_KEY> header. If you send an invalid key, you will receive a 401 error. If your prepaid credit is exhausted, you receive a 402 error. This ensures transparent usage without hidden fees.
Basic Completion Request
Send a standardPOST /v1/chat/completions request to generate text. Our uncensored AI API responds with text completions based on your prompt. The model ID is always uncensored. You can specify parameters like temperature and top_p to control randomness. The API accepts text input and returns text output. It does not support images, audio, or embeddings. Focus on text generation. You can set a max_tokens limit, with a default of 2,048 if not specified. The context window supports up to 64,000 tokens total (prompt + completion). Use this endpoint for straightforward text generation tasks.
Streaming Responses
Enable streaming by settingstream: true in your request. The API returns Server-Sent Events (SSE) chunks. Each chunk contains partial text. The final chunk includes token usage statistics. Streaming is ideal for real-time user experiences. It reduces perceived latency. Our uncensored AI models API handles streaming efficiently. You can parse chunks as they arrive. This allows instant display of generated text. Use streaming for chat interfaces or long-form content generation. It provides a smoother user experience compared to waiting for the full response. Token counts are only available in the last chunk.
Function Calling (Tools)
Support function calling for structured interactions. Define tools in your request with names, descriptions, and parameter schemas. The model can choose to call a function. Settool_choice to auto or specify a tool. This enables integration with external systems. Use this for data extraction or action triggers. The API supports standard OpenAI tool formats. It works with our uncensored model for reliable function selection. Define your tools clearly. The model returns a list of tool calls. Parse these calls to execute functions in your application. This adds capability beyond simple text generation.
JSON Mode
Force the model to output valid JSON. Setresponse_format: { "type": "json_object" } in your request. This ensures the output is parseable JSON. Useful for structured data extraction or API responses. The model adheres to JSON syntax. This reduces parsing errors in your code. Combine with function calling for robust structured outputs. Our uncensored AI API maintains this format reliably. Use it when you need consistent data structures. It simplifies downstream processing. Ensure your prompt encourages JSON output for best results.
Parameters, Limits, and Errors
Manage requests with specific limits. Max 300 requests per minute per key. Max 8 concurrent requests. Max 8 MB request body. Errors include 401 (invalid key), 402 (no credit), and 429 (rate limit). Context window is 64,000 tokens. Max output is 16,000 tokens. Use parameters liketemperature, top_p, stop, seed, presence_penalty, and frequency_penalty. These control generation behavior. Monitor your usage via token counts. Prepaid credit never expires. Errors are free. This predictable model suits stable integrations. Avoid exceeding limits to prevent 429 errors.cURL
curl https://api.unrestrictedaichatbot.cc/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'Python
from openai import OpenAI
client = OpenAI(base_url="https://api.unrestrictedaichatbot.cc/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Node.js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.unrestrictedaichatbot.cc/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Streaming
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Questions and answers
What is the maximum context window?
The context window supports up to 64,000 tokens, including both the prompt and the completion. This allows for substantial input data and long-form outputs within a single request.
How do I handle streaming responses?
Set the <code>stream</code> parameter to true in your request. The API returns Server-Sent Events (SSE) chunks containing partial text. Token usage is only provided in the final chunk.
Is there a concurrent request limit?
Yes, you can have up to 8 concurrent requests per API key. The rate limit is 300 requests per minute. Exceeding these limits will result in a 429 error.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.