Dokyumi Newsroom

Press Release

Dokyumi Releases Automated SIC Code Classification for Lenders, With a Confidence Score on Every Result

The classifier assigns both the 2-digit major group and the 4-digit industry code at application time, and flags low-confidence results for human review instead of guessing.

FOLSOM, Calif., August 22, 2026Dokyumi has released automated SIC code classification for lending and underwriting teams, assigning both the 2-digit major group and the specific 4-digit industry code from documents a lender already collects at application time, with a calibrated confidence score and a needs-review flag on every result.

Key figures and their sources

2-digit major group and 4-digit industry code
Both levels of granularity used in underwriting are returned on every classification, along with up to two plausible alternates and the evidence behind the choice.
Source: Dokyumi SIC classification documentation
filled application plus business bank statements
Classification runs against documents already collected at intake, so it adds no new document request to the application flow.
Source: Dokyumi SIC classification documentation
1987 SIC taxonomy, zero labeled examples
No labeled training set is required. The classifier works against the published taxonomy out of the box.
Source: Dokyumi SIC classification documentation
confidence score plus needs_review reason on 100% of results
Every result carries a confidence score and a needs-review flag with a stated reason, so a lender sets its own straight-through threshold rather than accepting an unqualified answer.
Source: Dokyumi SIC classification documentation

SIC coding at application time is usually either a manual underwriter judgment or a lookup against whatever the merchant typed into a free-text field. Both scale badly. The first consumes underwriter minutes on applications that will never fund, and the second inherits whatever the applicant decided to call their business.

Dokyumi reads the filled application and the business bank statements, which a lender already has, and resolves the industry code from the activity visible in both: payees, processors, suppliers and payroll patterns. Where the application and the bank statements disagree, that disagreement becomes the reason attached to a review flag rather than being silently resolved in favour of one of them.

The design choice worth naming is that the system is built to decline to answer. Thin signal, multi-line businesses, and application-versus-statement conflicts all route to a review queue with the reason spelled out. A classifier that returns a code for every application regardless of evidence is easy to build and hard to trust at volume.

Lenders can test the classifier against their own documents on the public demo page without an account, and validate accuracy on their own book by submitting a sample of recent applications and having underwriters spot-check the assigned codes.

Underwriters do not need a system that is right most of the time and silent about the rest. They need one that says which answers to trust, and hands back the ones it should not have guessed at.

John Arndt, Founder, Soxoa

About Dokyumi

Dokyumi extracts structured data from business documents including applications, bank statements and invoices, returning typed fields with confidence signals for review. It is operated by Soxoa. Dokyumi is not a credit bureau and does not provide underwriting or credit decisions.

Media contact: hello@dokyumi.com

Learn more: dokyumi.com/sic

Publisher: Dokyumi, operated by Soxoa

Published here only. Not distributed over a newswire.