Extract
Turn documents into structured JSON with per-field citations, in one call.
POST /v1/extract takes one or more documents and a JSON Schema. It returns data shaped like your schema. Add "citations": true to ground every field. Each field then points at the page and region it came from, or is declared absent.
How it works
- You send a file (a URL or a file id you already have) and an output shape. The shape is an inline
schemaor a savedaction. - CloudRaker fetches the bytes, parses the document, and runs extraction against your schema.
- The call holds until the run finishes, up to
?wait=seconds (60 by default, 120 max). If the run finishes in time, you get200with the full result. If not, you get202with a run id to poll. This is not an error. - The run and its files expire on their own (
ttl, 24 hours by default). Nothing accumulates in your organization.
Quickstart
The sample below uses a blank IRS Form W-9 as a public, stable test document. The call works with no local files. Replace the URL with your own when you are ready.
If you prefer a business document over a blank government form, download the sample invoice. It is one fictional page with a vendor, a bill-to, four line items, and totals. Send it with a presigned upload, then extract against its file id.
The TypeScript and Python samples are plain HTTP, so they run with nothing installed. The same call is one line on either SDK as of 0.3.0: client.extract({ file, schema }) in TypeScript, client.extract(file=…, schema=…) in Python. That page shows this exact extraction end to end in both languages.
Example response
Key fields
output appears only after status is processed.
Guaranteed grounding
With citations enabled, every filled field comes back cited or declared absent. Nothing is left ambiguous. A field the extractor filled without evidence gets a deterministic second model pass. That pass finds the citation or declares the value not present. A field declared not present is returned as null with a notFound citation, never a guess.
Cite against a recording
Extraction over audio grounds each field in the transcript, not in the audio timeline. An audio citation names the fileId, the quoted text, and a confidence. It carries no page, no bbox, and no timecode.
The transcript supplies the timing. Fetch it, then match the citation against it:
Two details make the match reliable. Join every segment into one normalized string, so a quote that crosses a segment boundary still resolves to where it begins. Then, if the whole quote does not match, retry with its first five to eight words — the platform trims long quotes at the tail, rarely at the head.
POST /v1/extract never inlines the transcript. Only GET /process/:id does, with include=results,content&format=json. On the verbs, plan for the extra file fetch.
Configuration
Every field below is optional unless noted.
processing on a file ref picks how the document is read: auto (default), ocr, simple, transcribe, or transcribe_diarize for audio.
One result over many documents
across_documents extracts each document independently. It then deterministically folds the per-document results into one record. It is not a joint pass over all files at once.
Per field, a real value always beats an empty one. The best-grounded value wins. With citations on, that is the value with the strongest citation. On ties, the earliest document in request order wins. With grounding off, every comparison is a tie.
Values are taken whole, together with their citations. Arrays are not unioned across documents. A list field comes from exactly one document.
Use it for related but independent documents that each contribute fields to one record. Examples: an application form plus a bank statement, a contract plus its amendment. It is not built for fragments of a single source. For a recording split into parts, concatenate the audio into one file before upload. Extraction then sees one continuous document.
Sync vs async
The endpoint is synchronous by default and degrades instead of failing.
When the run has not finished by the cap, you get 202 with a handle, never a timeout error:
Poll GET /v1/runs/:id (which also accepts ?wait=), or use a webhook. Send an idempotency-key header to make retries safe. A replay returns the original run and an idempotent-replay: true response header.
Large batches and scanned documents are the common causes of a 202. If you always want the handle, pass ?wait=0. Never block a request thread.
Schema inference
If you do not have a schema yet, omit both schema and action. CloudRaker infers a schema from the document, then extracts against it in the same call. Add hints (up to 2,000 characters) to steer what it looks for.
The finished run carries the schema it used at config.schema, alongside the output:
That response is trimmed. The real call on this document inferred 29 fields. The inferred schema is plain JSON Schema in the extraction dialect. Every field is nullable, and every field carries a description. You can send it back as schema with no editing.
config.schema is present whichever way the shape was decided: inferred, sent inline, or loaded from a saved config. One code path reads the applied shape.
Inference is for exploration, not production. The model picks the fields. Two runs over the same document can return different field names and a different field count. The two runs behind this page produced 32 and 29 fields, with form_number in one and form_type in the other. Nothing downstream of you can rely on that.
Use inference once to discover the shape, then pin it. Copy config.schema into your own request, or save it as an action and call that by name. Production callers must always send schema or action.
hints applies only to inference. Sending it alongside a schema or an action returns 400 invalid_request. Use instructions to guide an extraction whose shape you already fixed.
Save it as a config
Passing the same schema, instructions, and model on every call is repetitive. Save that configuration once in the extract config library. Then call {"file": …, "action": "invoice-lines"} and keep the request to two fields. Inline fields still win when you send them.
GET, PATCH and DELETE /v1/extract/configs/{idOrSlug} read, change and remove a config. GET /v1/extract/configs lists them. The flat /v1/actions routes still work as a deprecated alias over the same objects.
Citations are an add-on you enable per request or on the saved config. A saved config uses the setting’s internal name, grounding. It grounds its runs only when the config says "grounding": true, or when the request itself sends "citations": true.
The id and the slug are interchangeable wherever a config is referenced. Saved configs covers the catalog, the merge rules, and the benefits of the saved ramp. The same configurations are editable in the app under Actions.
Batch
POST /v1/extract/batch runs one saved config over many documents. It mints a separate run per file. Use it for a nightly backfill or a queue drain.
The response is always 202. Batches never run synchronously:
There is no batch object and no batch id. runs[] is the whole handle. Each entry is {id, statusUrl, file} for an accepted file. A file rejected at admission gets {file, error: {code, message}} instead. A bad URL in the list never sinks the rest.
Idempotency-Key is not honoured on a batch. One key cannot address N runs, so a retried batch fans out a second time. If a batch call fails ambiguously (a 429, a timeout), list its runs by shared metadata before you resend.
The body is strict. schema and hints are rejected with 400 invalid_request (“Unrecognized key”), not silently ignored. Save the schema as an action first. That is the point of the endpoint. A saved sign action passed as action is also a 400.
Track a batch
The batch costs one rate-limit token. Shared metadata is how you find its runs again:
That returns the three runs and their statuses in one call instead of three polls. See listing runs. For a per-run result, GET /v1/runs/:id stays authoritative. A webhook removes the polling entirely.
Next steps
The five rules your schema must satisfy, and the meaning of each rejection.
Mark a field as currency, a date, or a phone number for better reads.
Get clean markdown and structured JSON without a schema.
Save a schema once and reference it by id or slug.
Statuses, TTL, downloading outputs, and keeping a result.