Docs
Send a PDF or image to POST /v1/parse to parse it. Each page is processed independently, with its own status and charge. Get an API key from the dashboard and set it as OPEN_DOC_ROUTER_API_KEY, which the examples and SDKs read.
Every endpoint and field is in the API reference, generated from the OpenAPI spec.
Quickstart
curl https://www.opendocrouter.ai/v1/parse \
-H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemini-3.8-flash-low",
"document": { "url": "https://arxiv.org/pdf/1706.03762" }
}'Each example parses a public URL and gets every page back in the response. Each page has its own status, so check for failed pages before using the markdown. To send a local file, see Request; over 50 pages, see Large documents.
SDKs
The Python and TypeScript SDKs are typed clients for every endpoint, with retries and timeouts built in. They read your key from OPEN_DOC_ROUTER_API_KEY, or take it as api_key / apiKey.
| Language | Install | Source |
|---|---|---|
| Python | pip install opendocrouter | run-llama/opendocrouter-py |
| TypeScript | npm install @llamaindex/opendocrouter | run-llama/opendocrouter-ts |
Methods follow the endpoints: parse.create, parse.get, parse.delete, uploads.create, credits.get and models.list. In Python, a response's model_version is api_model_version, because Pydantic reserves the model_ prefix.
Models
Set the model to an ID from the models page or from GET /v1/models, such as google/gemini-3.8-flash-low.
Each model runs a parsing recipe (its prompt and settings), and the recipe has a version that changes whenever its output can change.
Request
POST /v1/parse with a JSON body and Authorization: Bearer <key>.
| Field | Description |
|---|---|
model | Required. A model ID. |
document | Required. One of the three forms below. |
pages | Optional, 1-based, like "1-3,7". Every page when left out. |
cache | Optional, default false. When true, the results are stored encrypted for 24 hours, readable with GET /v1/parse/<id>, and pages already stored for the same document and model come back free. Required for async requests. See Data retention. |
layout | Optional, default false. When true, each successful page also has a layout: where each part of its markdown is printed. See Layout. |
A document can be sent in three ways:
| Form | Description |
|---|---|
{ "url": "https://..." } | A public HTTPS URL, up to 50 MB or 500 pages. |
{ "upload_id": "..." } | A file you uploaded, up to 50 MB or 500 pages. |
{ "data": "<base64>", "mime_type": "application/pdf" } | Inline, up to about 3 MB. mime_type can also be image/png or image/jpeg. |
To upload, POST /v1/uploads returns a URL to PUT the file to within the hour. The file is deleted once a parse request reads it.
curl -X POST https://www.opendocrouter.ai/v1/uploads -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY"
# { "upload_id": "9b2c...", "upload_url": "https://...", "max_bytes": 52428800, ... }
curl -T report.pdf "<upload_url>"
# Then send { "document": { "upload_id": "9b2c..." } } to /v1/parseResponse
By default a request is sync (mode: "sync"), for up to 50 pages, and the response is a 200 with every page. A status of completed means every page worked, partial means some did, and failed means none did. Join pages[].markdown for a single markdown string.
{
"id": "0f433d25-c2d5-40df-960f-1cb4c5ddc415",
"status": "partial",
"model": "google/gemini-3.8-flash-low",
"model_version": "2026-09-28",
"price_version": "2026-09-25",
"pages": [
{ "page": 1, "status": "ok", "markdown": "# Attention Is All You Need ...", "cached": false,
"usage": { "input_tokens": 757, "output_tokens": 811 }, "charge_usd": 0.00397 },
{ "page": 2, "status": "error",
"error": { "code": "timeout", "message": "The page didn't finish before the request deadline" },
"charge_usd": 0 }
],
"usage": { "input_tokens": 757, "output_tokens": 811 },
"charge_usd": 0.00397
}GET /v1/parse/<id> shows what any request did and cost, page by page. Without cache: true it doesn't have a sync request's markdown, which is only in the POST response: if your connection drops, the request is still charged, and sending it again is a new request. With cache: true, add expand=markdown to read it again, and sending the same request again is free.
Layout
With layout: true, each successful page gets a layout: the page's size, and its elements in reading order. An element has a type (title, section_header, text, list_item, table, picture, chart, formula, caption, footnote, page_header, page_footer, code, form or key_value), the lines of the page's markdown it spans, and its boxes on the page.
{
"page": 1, "status": "ok", "markdown": "# Attention Is All You Need\n\nThe dominant sequence ...",
"layout": {
"status": "ok", "width": 612, "height": 792,
"elements": [
{ "type": "title", "lines": [0, 0], "confidence": 0.94,
"boxes": [{ "x": 0.27, "y": 0.11, "w": 0.46, "h": 0.03 }] },
{ "type": "text", "lines": [2, 2], "confidence": 0.91,
"boxes": [{ "x": 0.12, "y": 0.18, "w": 0.36, "h": 0.32 }, { "x": 0.52, "y": 0.18, "w": 0.36, "h": 0.12 }] },
{ "type": "picture", "lines": null, "confidence": 0.82,
"boxes": [{ "x": 0.55, "y": 0.34, "w": 0.3, "h": 0.2 }] }
]
}
}linesis the first and last line, counting from 0, after splitting the markdown on"\n"alone. Python'ssplitlines()also splits on other characters, so usesplit("\n"). It'snullfor a picture on the page that the markdown doesn't mention.- A box's
x,y,wandhare fractions of the page, from its top left.r, when present, is the clockwise angle of text printed at a slant. An element printed in pieces, such as a paragraph that continues in the next column, has a box per piece; one that couldn't be placed has none. - If the layout of a page fails, the page keeps its markdown and its
layouthas astatusoferrorwith acodeandmessage. The layout isn't charged.
Layout adds $0.20 per million tokens to each page whose layout comes back. For async requests, add expand=layout (or expand=markdown,layout) to the poll.
Large documents
Over 50 pages, send "mode": "async" with "cache": true (up to 500 pages; a sync request over 50 is refused with too_large). Async requests store their results, so they need cache. The request returns a 202 with status: "processing" and an id. Poll GET /v1/parse/<id> until the status changes; pages_done shows progress. Then add expand=markdown to get the markdown. Every page usually comes back at once; results over 4 MB come in parts: while has_more is true, pass next_cursor as cursor.
curl https://www.opendocrouter.ai/v1/parse \
-H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemini-3.8-flash-low",
"document": { "url": "https://arxiv.org/pdf/1706.03762" },
"mode": "async",
"cache": true
}'
# { "id": "0f433d25-...", "status": "processing", ... }
# Poll until status isn't "processing"
curl "https://www.opendocrouter.ai/v1/parse/<id>" -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY"
# Then fetch the markdown
curl "https://www.opendocrouter.ai/v1/parse/<id>?expand=markdown" -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY"
# Only if has_more is true (results over 4 MB): fetch the rest
curl "https://www.opendocrouter.ai/v1/parse/<id>?expand=markdown&cursor=<next_cursor>" -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY"The markdown is kept encrypted for 24 hours, until results_expire_at. DELETE /v1/parse/<id> deletes it sooner, and stops a running request; only pages that already ran are charged. Either way, the request's status and cost stay available.
Data retention
Nothing from your document is kept unless you send cache: true:
- Without
cache(sync only): the markdown is only in the POST response. The document is never stored. - With
cache: true(required for async): the results are kept encrypted for 24 hours after the request finishes, untilresults_expire_at.GET /v1/parse/<id>?expand=markdownreads them, and a later request for the same pages of the same document, with the same model andlayout, is served from them for free. An async request's document is kept only until parsing ends.
DELETE /v1/parse/<id> deletes a request's stored results sooner, and they're no longer served from the cache. Each request's status, page count and cost stay in your history.
Retrying pages
We retry each page once on timeouts, rate limits, provider errors and capacity, and when the model loops or doesn't transcribe. If a page still fails, send the document again with just those pages or try a different model. An uploaded file only works once, so upload the file again or use a URL.
{
"model": "google/gemini-3.8-flash-low",
"document": { "url": "https://arxiv.org/pdf/1706.03762" },
"pages": "2,7"
}A page error has a code, a message, and a reason when the provider gave one (such as SAFETY or max_tokens). From GET /v1/parse/<id>, message is the code's standard description unless you ask for expand=markdown.
| Page error | Meaning | Worth retrying |
|---|---|---|
timeout | Still running at the deadline, or the provider timed out | Yes |
rate_limited | The provider rate-limited us | Yes |
provider_error | The provider returned an error; its message is included | Yes |
at_capacity | No capacity came free within 10 seconds | Yes |
output_truncated | The page hit the model's output limit | Rarely |
content_filtered | The provider blocked the output, or the model declined | Rarely |
repetitive_output | The model got stuck repeating text | Sometimes |
invalid_output | The model answered without transcribing the page | Sometimes |
empty_output | The model returned no text | Sometimes |
response_too_large | The page didn't fit in the 4.5 MB response | Yes, on its own |
unreadable_page | The PDF opened, but this page couldn't be read | No |
not_processed | The async request ended (it failed or was deleted) before this page ran | Yes |
Errors
A request that can't run returns { "error": { "code", "message" } } and is never charged. Every response has an X-Request-Id header, which also appears in your usage log.
| Status | Code | When |
|---|---|---|
| 400 | invalid_request | Bad JSON, unknown model, or a bad pages value |
| 400 | url_not_allowed | The URL isn't public HTTPS on the default port, or redirects somewhere that isn't |
| 401 | unauthorized | Missing, unknown or revoked API key |
| 402 | insufficient_credits | Not enough credit to cover the request's maximum charge. Includes required_usd and available_usd |
| 403 | account_paused | The account is paused; email support@runllama.ai |
| 404 | not_found | No such request or upload on this account, or the upload was used |
| 409 | results_not_stored | expand on a sync request sent without cache: true, whose markdown and layout are only in the POST response |
| 410 | gone | expand after a request's stored results were deleted, by you or 24 hours after it finished |
| 413 | too_large | A sync request over 50 pages, any request over 500, a request body over 4 MB, or a file over 50 MB |
| 415 | unsupported_type | Not a PDF, PNG or JPEG |
| 422 | unreadable_document | An encrypted or corrupt PDF, or a URL that couldn't be fetched |
| 429 | rate_limited | Over the account's concurrency limits, or the provider rate-limited every page. Includes Retry-After |
| 503 | at_capacity | The model is at capacity or paused. Includes Retry-After |
| 503 | model_starting | A model we host is starting up. Includes Retry-After (a few minutes) |
Billing
Credit is prepaid: new accounts start with $5 of free credit, and you top up from $10 in the dashboard, plus a 5% fee on each top-up. Frontier models are charged at their providers' token prices, with no markup. Each successful page is charged for its tokens at the model's price, plus the layout price when you ask for layout. Failed, cached and blank pages are free.
To start, a request needs enough credit for its maximum charge: the model's most per page, times the number of pages. That amount is held while it runs, and whatever the pages didn't use is released when it finishes.
Limits
| Limit | Value |
|---|---|
| Pages | Up to 50 per sync request, 500 per async request. |
| Inline files | About 3 MB (the request body is capped at 4 MB). |
| URLs and uploads | Up to 50 MB. |
| Concurrency | 10 requests, and 5 asynchronous requests, running at once per account. |
| Requests | 300 POST /v1/parse calls a minute per account. |
| Polling | 60 GET /v1/parse/{id} calls a minute per account. |
| Time | Pages still running after 270 seconds come back as timeout. Asynchronous requests wait up to 30 minutes for capacity. |