PDF Data Extraction

中级data最低上下文:32K

Extracts structured data from PDFs and scanned documents — invoices, receipts, forms, contracts, reports, and tables. Returns clean, typed output (JSON, CSV, or Markdown tables), handles multi-page layouts and nested tables, and flags low-confidence fields for review. Uses vision-capable models for image-based and scanned PDFs.

使用场景

  • Extracting line items and totals from invoices and receipts
  • Converting PDF tables into clean CSV or JSON
  • Pulling structured fields from forms and applications
  • Digitizing scanned documents into machine-readable data
  • Extracting key terms and clauses from contracts

示例提示词

Extract structured data from the attached invoice PDF. Return JSON with:
- vendor name, invoice number, issue date, due date
- line items (description, quantity, unit price, amount)
- subtotal, tax, and total
- currency

Flag any field you are less than 90% confident about under a "needs_review" key.

推荐模型

兼容工具

claude-codekiroany

模态

输入: file, image, text
输出: text, file

相关 Skills

作者

OpenModels Community

@openmodelsrun