Worked implementation guide · reviewed September 16, 2026

Custom schema extraction: from field definition to API response

Dokyumi’s extraction request does not send a JSON schema inline. Create the schema first, then send the document and that schema’s slug to the shared extraction endpoint.

Direct answer

Define named, typed fields in the dashboard; POST a supported file plus the schema slug to/api/v1/extract; then branch on status and inspect both validation arrays before using the returned data.

Run the review gate before connecting a destination

A missing required value fails validation and puts the extraction in review. This worked invoice example returns a review status and a missing total. Type validation still does not prove that a printed total reconciles with the source. The small Node.js script holds it for review. It does not call an API, upload a document, or spend credits.

node review-extraction.mjs invoice-review-response.json
# decision: hold_for_review
# reason: total_amount is missing or invalid for this invoice workflow.

Changing the total alone keeps this example in review: its status and validation errors still require attention. Only a completed response with valid fields and no validation errors can become ready_for_source_comparison. Payment approval stays false. This illustrates your application’s review policy, not an OCR result or an accuracy claim. Adapt the fields, credit-note rules, dates, and destination checks to your own process.

Evaluate the whole handoff

  1. Choose documents that represent your layouts, page counts, and source quality. Keep a reviewer’s reference values.
  2. Create the schema below, run authenticated extraction, and save the returned status, validation details, and credits used.
  3. Count corrected required fields, held records, and review minutes. Confirm destination mappings before enabling writes.

Example and handling guidance updated September 16, 2026. The example above is illustrative; the measured synthetic API pilot appears below.

Choose a workflow-specific starter schema

Step 1

Define the extraction schema

This example mirrors Dokyumi’s canonical field shape. Create it through the dashboard; the public extraction endpoint accepts the resulting slug, not this object as its request body. Supported field types include string, number, currency, date, boolean, enum, array, and object.

Resulting schema definition · invoice-intake
{
  "name": "Invoice Intake",
  "slug": "invoice-intake",
  "description": "Fields used by the AP intake workflow",
  "ocr_mode": "standard",
  "confidence_threshold": 0.8,
  "fields": [
    {
      "key": "vendor_name",
      "label": "Vendor Name",
      "type": "string",
      "required": true,
      "description": "Company or person issuing the invoice"
    },
    {
      "key": "invoice_number",
      "label": "Invoice Number",
      "type": "string",
      "required": true
    },
    {
      "key": "invoice_date",
      "label": "Invoice Date",
      "type": "date",
      "required": true,
      "description": "ISO 8601 date: YYYY-MM-DD"
    },
    {
      "key": "purchase_order_number",
      "label": "Purchase Order Number",
      "type": "string",
      "required": false
    },
    {
      "key": "total_amount",
      "label": "Total Amount",
      "type": "currency",
      "required": true,
      "description": "Final amount due as a number"
    }
  ]
}
The dashboard creation route derives the slug from the schema name. Field keys allow letters, numbers, and underscores; duplicate keys are rejected. The current dashboard creation flow uses the default confidence threshold of 0.8; the public extraction call does not accept a per-request threshold.

Step 2

Send the file and schema slug

The endpoint is synchronous and uses Bearer authentication plus multipart form data. Althoughschema is optional at the protocol level, passing it explicitly avoids the fallback to your first active schema. Direct API calls do not accept a webhook URL; configurable webhook delivery belongs to separate white-label upload sites.

cURL requestPOST /api/v1/extract
curl -X POST https://dokyumi.com/api/v1/extract \
  -H "Authorization: Bearer dk_live_your_api_key" \
  -F "file=@invoice.pdf" \
  -F "schema=invoice-intake"
FieldTypeWhereMeaning
filemultipart fileRequiredA supported PDF or image file, up to 20MB.
schemastringOptional in the protocolThe active schema slug. Pass it explicitly so the request cannot fall back to the organization’s first active schema.
statuscompleted | reviewSuccess responseBranch first. Failures use a non-2xx error envelope, not a failed success status.
dataobjectSuccess responseModel-produced fields. Treat them as accepted only after your validation and review policy.
confidencepartial field mapSuccess responseModel-reported 0–1 values when present; fields can be omitted.
validationobjectSuccess responseIncludes valid, errors, and low_confidence_fields. Inspect both arrays for review results.
metaobjectSuccess responseIncludes processing time, page count, credits used, OCR-cache state, and runtime model identifier.

Step 3

Read the success envelope

The template below assumes a one-page example that passed type validation. Angle-bracket values are placeholders, not literal API values; no timing or model performance is implied. The empty confidence map is valid because reported field scores are partial and may be omitted.

Annotated response template · placeholders are non-literal
{
  "id": "<extraction UUID>",
  "status": "completed",
  "request_id": "<request UUID>",
  "schema": "invoice-intake",
  "data": {
    "vendor_name": "Acme Corp",
    "invoice_number": "INV-2026-001",
    "invoice_date": "2026-08-20",
    "purchase_order_number": null,
    "total_amount": 1250
  },
  "confidence": {},
  "validation": {
    "valid": true,
    "errors": [],
    "low_confidence_fields": []
  },
  "meta": {
    "processing_time_ms": "<measured integer>",
    "page_count": 1,
    "credits_used": 1,
    "ocr_cached": false,
    "model": "<runtime model identifier>"
  }
}

completed

No Zod type-validation errors and no reported confidence values below the schema threshold. It is not a guarantee that every value is correct.

review

At least one validation error or reported below-threshold score. Inspect validation.errors and validation.low_confidence_fields.

non-2xx

The request failed. Read error and code; page-limit errors can also include pages and page_limit.

Step 4

Make review handling part of the integration

Handle the HTTP boundary first, then route successful review results using both sources of review detail. Your business rules can add stricter checks for nulls, required identifiers, totals, or cross-field consistency before writing anywhere downstream.

Server-side JavaScript pattern
const response = await fetch('https://dokyumi.com/api/v1/extract', {
  method: 'POST',
  headers: { Authorization: `Bearer ${process.env.DOKYUMI_API_KEY}` },
  body: formData,
})

if (!response.ok) {
  const failure = await response.json()
  throw new Error(`${failure.code}: ${failure.error}`)
}

const result = await response.json()

if (result.status !== 'completed' ||
    result.validation?.valid !== true ||
    result.validation.errors.length ||
    result.validation.low_confidence_fields.length ||
    result.data?.total_amount == null) {
  await sendToReviewQueue({
    extractionId: result.id,
    data: result.data,
    validationErrors: result.validation.errors,
    lowConfidenceFields: result.validation.low_confidence_fields,
  })
} else {
  await compareFieldsWithSource(result.data) // Your application owns this step.
}

HTTP 400

Missing/invalid file or no active matching schema.

HTTP 402

The post-OCR page count requires more credits than remain.

HTTP 403

The API key scope does not authorize that schema.

HTTP 413

The document exceeds the configured page limit; self-serve plans use 50.

HTTP 429

The configured request limiter or monthly credit quota was reached.

Production checklist

What to verify on your documents

  • Use representative digital PDFs, scans, photos, and layout variants from the workflow you intend to automate.
  • Confirm that field types match downstream expectations; currency and number fields validate as numbers, while date fields use YYYY-MM-DD.
  • Treat null required values and omitted confidence scores according to explicit business rules.
  • Estimate credits with ceil(page_count / 5); one extraction can use more than one credit.
  • Keep API keys server-side and scope keys when a caller should access only selected schemas.
  • Review the security page before using documents with sensitive data, and document your deletion-request process.

A measured custom-schema pilot

On September 16, 2026 Pacific time, we sent three synthetic one-page invoices to the production API, then repeated them after fixing required-field validation. The first run exposed a bug: a missing required total was returned as null but marked valid and completed. The retest correctly returned review with a validation error, while preserving the missing value.

Three production test cases before and after the required-field validation fix
Synthetic inputReturned total, both runsBeforeAfter
Clean invoice PDF250completed · validcompleted · valid
Mismatched total PDF300completed · validcompleted · valid
Missing total PDFnullcompleted · validreview · invalid

The mismatched invoice prints a $300 total against $250 of line items. This schema extracts header fields, so a valid number does not establish that its arithmetic is correct. Compare line items and source values before approving a downstream write. The absent optional purchase order stayed null and valid in both runs.

All six requests returned HTTP 200 and used one existing QA credit each. Initial OCR was uncached; retest OCR was cached, so timings are not a controlled speed comparison. Temporary scoped keys were revoked and schemas deactivated. This small synthetic API pilot does not establish scan accuracy, customer outcomes, or signup and checkout behavior. Model-reported confidence is not a substitute for validation.

Inspect the schema and before/after responses

Evidence notes

Sources and limitations

Sources used

Limitations

  • The worked response is illustrative; the separately labeled pilot reports actual synthetic API requests. Neither is a customer case or a representative accuracy study.
  • The response block uses explicit placeholders for request-specific timing and model values and therefore is an annotated template, not literal JSON to paste into a parser.
  • Model-reported confidence is not calibrated accuracy, and an omitted confidence value is not approval.
  • Type validation does not establish source-document truth. Add business checks and human review appropriate to the consequences of a wrong field.

Last reviewed September 16, 2026. Product behavior can change; the linked API, pricing, and security pages are the controlling public references.