Documents
Extract text from uploaded documents (CVs, certificates): PDF with per-page OCR fallback, DOCX, plain text and images. Runs on our own servers; no document is sent to a third party. Members ingest their own gated-files upload, read back only their own result, and confirm or reject the skill candidates found in it.
Endpoints (9)
Queue text extraction for your own member upload. Requires a verified tenant_auth session (wst_ Bearer). The file must be a gated-files member upload by you, in your tenant. Idempotent: offering the same file again returns the same extraction. Returns { extraction_id, status, record_id, existing }.
| Field | Type | Required | Description |
|---|---|---|---|
file_id |
integer | ✓ Yes | file_id from gated-files/put-upload. |
Status and result of one of your own extractions. Requires a verified tenant_auth session. Returns status (queued | extracting | ok | partial | failed), method, quality, errors, notes and per-page metadata; the text only with include_text=true.
| Field | Type | Required | Description |
|---|---|---|---|
extraction_id |
integer | ✓ Yes | extraction_id from ingest. |
include_text |
boolean | No | Include the extracted text (default false). |
Skill candidates from one of your own extractions, for review. Requires a verified tenant_auth session. Returns candidates_status (queued | running | ok | failed; null before the first run) and the candidates: occupations, software and skills, each grounded in O*NET with the CV lines as evidence. Default filter: pending. After a new candidate run, undecided candidates get new candidate_ids — fetch the list again before deciding.
| Field | Type | Required | Description |
|---|---|---|---|
extraction_id |
integer | ✓ Yes | extraction_id from ingest. |
status |
string | No | Filter: pending (default) | confirmed | rejected | all. |
Confirm or reject skill candidates from one of your own extractions. Requires a verified tenant_auth session and skill confirmation enabled for the project. A confirmed candidate becomes a record in the configured skill entity under your identity, with its computation trace; a rejected candidate stays out of your profile. A decision is final: deciding again returns skipped, an unknown or outdated candidate_id returns not_found. Returns { extraction_id, results, counts }.
| Field | Type | Required | Description |
|---|---|---|---|
extraction_id |
integer | ✓ Yes | extraction_id from ingest. |
decisions |
array | ✓ Yes | List of { candidate_id, decision: confirm | reject }; at most 200, each candidate_id once. |
Project settings. ingest_enabled switches member ingest on (default off). record_entity names the MAPI entity for the per-upload anchor record; it must have the fields gated_file_id, extraction_id and original_filename, an owner_field and create: { verified: self }. skill_entity names the MAPI entity for confirmed candidates; it must have the trace fields item_kind, onet_ref, onet_category, match_type, match_score, evidence, candidate_id, extraction_id, method_version and onet_release, an owner_field and create: { verified: self }. skill_field_map maps anchor_record_id, label and rank to field names of that entity. All checked on save. Send an empty record_entity to switch the anchor record off, an empty skill_entity to switch confirmation off.
| Field | Type | Required | Description |
|---|---|---|---|
ingest_enabled |
boolean | No | Allow members to ingest uploads (default false). |
record_entity |
string | No | MAPI entity for the anchor record, or empty for none. |
skill_entity |
string | No | MAPI entity for confirmed candidates, or empty to switch confirmation off. |
skill_field_map |
object | No | Field names in skill_entity: { anchor_record_id, label, rank }. |
Full extraction for the owner: status, text, per-page text and methods, quality, engines, errors and notes. For review and for the skill-extraction step.
| Field | Type | Required | Description |
|---|---|---|---|
extraction_id |
integer | ✓ Yes | Extraction ID. |
Which extraction programs are installed on the serving node, with versions and OCR languages. Diagnostics after an install or deploy.
No input parameters required.
Skill candidates found in one extraction: occupations (from job titles), software and skills, each grounded in the occupations reference, with the CV lines as evidence, the method version and the O*NET release. Candidates are not profile data until the member confirms them.
| Field | Type | Required | Description |
|---|---|---|---|
extraction_id |
integer | ✓ Yes | Extraction ID. |
status |
string | No | Filter: pending | confirmed | rejected. Default all. |
Queue a new candidate run for an extraction, for example after the method improved. Candidates the member already confirmed or rejected stay as they are.
| Field | Type | Required | Description |
|---|---|---|---|
extraction_id |
integer | ✓ Yes | Extraction ID. |
MCP Tool Names
When using this integration through an AI assistant (Claude, ChatGPT, Cursor, etc.), the endpoints are available as MCP tools:
| Endpoint | MCP Tool Name |
|---|---|
| ingest | documents_ingest |
| status | documents_status |
| my-candidates | documents_my_candidates |
| decide | documents_decide |
| configure | documents_configure |
| extraction | documents_extraction |
| tools-status | documents_tools_status |
| candidates | documents_candidates |
| regenerate | documents_regenerate |
Website