Skip to content

Docs

Send a PDF or image to POST /v1/parse to parse it. Each page is processed independently, with its own status and charge. Get an API key from the dashboard and set it as OPEN_DOC_ROUTER_API_KEY, which the examples and SDKs read.

Every endpoint and field is in the API reference, generated from the OpenAPI spec.

Quickstart

curl https://www.opendocrouter.ai/v1/parse \
  -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemini-3.8-flash-low",
    "document": { "url": "https://arxiv.org/pdf/1706.03762" }
  }'

Each example parses a public URL and gets every page back in the response. Each page has its own status, so check for failed pages before using the markdown. To send a local file, see Request; over 50 pages, see Large documents.

SDKs

The Python and TypeScript SDKs are typed clients for every endpoint, with retries and timeouts built in. They read your key from OPEN_DOC_ROUTER_API_KEY, or take it as api_key / apiKey.

LanguageInstallSource
Pythonpip install opendocrouterrun-llama/opendocrouter-py
TypeScriptnpm install @llamaindex/opendocrouterrun-llama/opendocrouter-ts

Methods follow the endpoints: parse.create, parse.get, parse.delete, uploads.create, credits.get and models.list. In Python, a response's model_version is api_model_version, because Pydantic reserves the model_ prefix.

Models

Set the model to an ID from the models page or from GET /v1/models, such as google/gemini-3.8-flash-low.

Each model runs a parsing recipe (its prompt and settings), and the recipe has a version that changes whenever its output can change.

Request

POST /v1/parse with a JSON body and Authorization: Bearer <key>.

FieldDescription
modelRequired. A model ID.
documentRequired. One of the three forms below.
pagesOptional, 1-based, like "1-3,7". Every page when left out.
cacheOptional, default false. When true, the results are stored encrypted for 24 hours, readable with GET /v1/parse/<id>, and pages already stored for the same document and model come back free. Required for async requests. See Data retention.
layoutOptional, default false. When true, each successful page also has a layout: where each part of its markdown is printed. See Layout.

A document can be sent in three ways:

FormDescription
{ "url": "https://..." }A public HTTPS URL, up to 50 MB or 500 pages.
{ "upload_id": "..." }A file you uploaded, up to 50 MB or 500 pages.
{ "data": "<base64>", "mime_type": "application/pdf" }Inline, up to about 3 MB. mime_type can also be image/png or image/jpeg.

To upload, POST /v1/uploads returns a URL to PUT the file to within the hour. The file is deleted once a parse request reads it.

curl -X POST https://www.opendocrouter.ai/v1/uploads -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY"
# { "upload_id": "9b2c...", "upload_url": "https://...", "max_bytes": 52428800, ... }

curl -T report.pdf "<upload_url>"

# Then send { "document": { "upload_id": "9b2c..." } } to /v1/parse

Response

By default a request is sync (mode: "sync"), for up to 50 pages, and the response is a 200 with every page. A status of completed means every page worked, partial means some did, and failed means none did. Join pages[].markdown for a single markdown string.

{
  "id": "0f433d25-c2d5-40df-960f-1cb4c5ddc415",
  "status": "partial",
  "model": "google/gemini-3.8-flash-low",
  "model_version": "2026-09-28",
  "price_version": "2026-09-25",
  "pages": [
    { "page": 1, "status": "ok", "markdown": "# Attention Is All You Need ...", "cached": false,
      "usage": { "input_tokens": 757, "output_tokens": 811 }, "charge_usd": 0.00397 },
    { "page": 2, "status": "error",
      "error": { "code": "timeout", "message": "The page didn't finish before the request deadline" },
      "charge_usd": 0 }
  ],
  "usage": { "input_tokens": 757, "output_tokens": 811 },
  "charge_usd": 0.00397
}

GET /v1/parse/<id> shows what any request did and cost, page by page. Without cache: true it doesn't have a sync request's markdown, which is only in the POST response: if your connection drops, the request is still charged, and sending it again is a new request. With cache: true, add expand=markdown to read it again, and sending the same request again is free.

Layout

With layout: true, each successful page gets a layout: the page's size, and its elements in reading order. An element has a type (title, section_header, text, list_item, table, picture, chart, formula, caption, footnote, page_header, page_footer, code, form or key_value), the lines of the page's markdown it spans, and its boxes on the page.

{
  "page": 1, "status": "ok", "markdown": "# Attention Is All You Need\n\nThe dominant sequence ...",
  "layout": {
    "status": "ok", "width": 612, "height": 792,
    "elements": [
      { "type": "title", "lines": [0, 0], "confidence": 0.94,
        "boxes": [{ "x": 0.27, "y": 0.11, "w": 0.46, "h": 0.03 }] },
      { "type": "text", "lines": [2, 2], "confidence": 0.91,
        "boxes": [{ "x": 0.12, "y": 0.18, "w": 0.36, "h": 0.32 }, { "x": 0.52, "y": 0.18, "w": 0.36, "h": 0.12 }] },
      { "type": "picture", "lines": null, "confidence": 0.82,
        "boxes": [{ "x": 0.55, "y": 0.34, "w": 0.3, "h": 0.2 }] }
    ]
  }
}
  • lines is the first and last line, counting from 0, after splitting the markdown on "\n" alone. Python's splitlines() also splits on other characters, so use split("\n"). It's null for a picture on the page that the markdown doesn't mention.
  • A box's x, y, w and h are fractions of the page, from its top left. r, when present, is the clockwise angle of text printed at a slant. An element printed in pieces, such as a paragraph that continues in the next column, has a box per piece; one that couldn't be placed has none.
  • If the layout of a page fails, the page keeps its markdown and its layout has a status of error with a code and message. The layout isn't charged.

Layout adds $0.20 per million tokens to each page whose layout comes back. For async requests, add expand=layout (or expand=markdown,layout) to the poll.

Large documents

Over 50 pages, send "mode": "async" with "cache": true (up to 500 pages; a sync request over 50 is refused with too_large). Async requests store their results, so they need cache. The request returns a 202 with status: "processing" and an id. Poll GET /v1/parse/<id> until the status changes; pages_done shows progress. Then add expand=markdown to get the markdown. Every page usually comes back at once; results over 4 MB come in parts: while has_more is true, pass next_cursor as cursor.

curl https://www.opendocrouter.ai/v1/parse \
  -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemini-3.8-flash-low",
    "document": { "url": "https://arxiv.org/pdf/1706.03762" },
    "mode": "async",
    "cache": true
  }'
# { "id": "0f433d25-...", "status": "processing", ... }

# Poll until status isn't "processing"
curl "https://www.opendocrouter.ai/v1/parse/<id>" -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY"

# Then fetch the markdown
curl "https://www.opendocrouter.ai/v1/parse/<id>?expand=markdown" -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY"

# Only if has_more is true (results over 4 MB): fetch the rest
curl "https://www.opendocrouter.ai/v1/parse/<id>?expand=markdown&cursor=<next_cursor>" -H "Authorization: Bearer $OPEN_DOC_ROUTER_API_KEY"

The markdown is kept encrypted for 24 hours, until results_expire_at. DELETE /v1/parse/<id> deletes it sooner, and stops a running request; only pages that already ran are charged. Either way, the request's status and cost stay available.

Data retention

Nothing from your document is kept unless you send cache: true:

  • Without cache (sync only): the markdown is only in the POST response. The document is never stored.
  • With cache: true (required for async): the results are kept encrypted for 24 hours after the request finishes, until results_expire_at. GET /v1/parse/<id>?expand=markdown reads them, and a later request for the same pages of the same document, with the same model and layout, is served from them for free. An async request's document is kept only until parsing ends.

DELETE /v1/parse/<id> deletes a request's stored results sooner, and they're no longer served from the cache. Each request's status, page count and cost stay in your history.

Retrying pages

We retry each page once on timeouts, rate limits, provider errors and capacity, and when the model loops or doesn't transcribe. If a page still fails, send the document again with just those pages or try a different model. An uploaded file only works once, so upload the file again or use a URL.

{
  "model": "google/gemini-3.8-flash-low",
  "document": { "url": "https://arxiv.org/pdf/1706.03762" },
  "pages": "2,7"
}

A page error has a code, a message, and a reason when the provider gave one (such as SAFETY or max_tokens). From GET /v1/parse/<id>, message is the code's standard description unless you ask for expand=markdown.

Page errorMeaningWorth retrying
timeoutStill running at the deadline, or the provider timed outYes
rate_limitedThe provider rate-limited usYes
provider_errorThe provider returned an error; its message is includedYes
at_capacityNo capacity came free within 10 secondsYes
output_truncatedThe page hit the model's output limitRarely
content_filteredThe provider blocked the output, or the model declinedRarely
repetitive_outputThe model got stuck repeating textSometimes
invalid_outputThe model answered without transcribing the pageSometimes
empty_outputThe model returned no textSometimes
response_too_largeThe page didn't fit in the 4.5 MB responseYes, on its own
unreadable_pageThe PDF opened, but this page couldn't be readNo
not_processedThe async request ended (it failed or was deleted) before this page ranYes

Errors

A request that can't run returns { "error": { "code", "message" } } and is never charged. Every response has an X-Request-Id header, which also appears in your usage log.

StatusCodeWhen
400invalid_requestBad JSON, unknown model, or a bad pages value
400url_not_allowedThe URL isn't public HTTPS on the default port, or redirects somewhere that isn't
401unauthorizedMissing, unknown or revoked API key
402insufficient_creditsNot enough credit to cover the request's maximum charge. Includes required_usd and available_usd
403account_pausedThe account is paused; email support@runllama.ai
404not_foundNo such request or upload on this account, or the upload was used
409results_not_storedexpand on a sync request sent without cache: true, whose markdown and layout are only in the POST response
410goneexpand after a request's stored results were deleted, by you or 24 hours after it finished
413too_largeA sync request over 50 pages, any request over 500, a request body over 4 MB, or a file over 50 MB
415unsupported_typeNot a PDF, PNG or JPEG
422unreadable_documentAn encrypted or corrupt PDF, or a URL that couldn't be fetched
429rate_limitedOver the account's concurrency limits, or the provider rate-limited every page. Includes Retry-After
503at_capacityThe model is at capacity or paused. Includes Retry-After
503model_startingA model we host is starting up. Includes Retry-After (a few minutes)

Billing

Credit is prepaid: new accounts start with $5 of free credit, and you top up from $10 in the dashboard, plus a 5% fee on each top-up. Frontier models are charged at their providers' token prices, with no markup. Each successful page is charged for its tokens at the model's price, plus the layout price when you ask for layout. Failed, cached and blank pages are free.

To start, a request needs enough credit for its maximum charge: the model's most per page, times the number of pages. That amount is held while it runs, and whatever the pages didn't use is released when it finishes.

Limits

LimitValue
PagesUp to 50 per sync request, 500 per async request.
Inline filesAbout 3 MB (the request body is capped at 4 MB).
URLs and uploadsUp to 50 MB.
Concurrency10 requests, and 5 asynchronous requests, running at once per account.
Requests300 POST /v1/parse calls a minute per account.
Polling60 GET /v1/parse/{id} calls a minute per account.
TimePages still running after 270 seconds come back as timeout. Asynchronous requests wait up to 30 minutes for capacity.