> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bono.network/llms.txt
> Use this file to discover all available pages before exploring further.

# OCR & Scanned Documents

> Process scanned PDFs, images, and documents with Optical Character Recognition

## Overview

Arbiter's OCR (Optical Character Recognition) processing extracts text from scanned documents, images, and PDFs that don't have selectable text. This enables full analysis and AI features on documents that would otherwise be unreadable.

<Info>
  **OCR Pricing:** 5 tokens per page (compared to 1 token per page for standard text-based documents).
</Info>

## When to Use OCR

Use OCR processing for:

* **Scanned PDFs** - Documents scanned from paper
* **Image files** - Photos of documents (PNG, JPG, etc.)
* **Protected PDFs** - Some PDFs with restricted text selection
* **Old documents** - Archived materials in image format

### Automatic Detection

Arbiter automatically detects if a PDF likely needs OCR by analyzing:

* Text density per page
* Average words extracted
* Image-to-text ratio

When a document appears to be scanned (\< 20 words per page detected), Arbiter:

1. Alerts you that OCR may be needed
2. Automatically enables OCR option
3. Shows updated cost estimate

## Uploading Documents with OCR

### Single Document Upload

<Steps>
  <Step title="Click Upload Document">
    From dashboard or sidebar
  </Step>

  <Step title="Select Your File">
    Choose your PDF, image, or document
  </Step>

  <Step title="Review OCR Detection">
    Arbiter shows if OCR is recommended:

    * "Likely scanned document detected"
    * "OCR recommended for best results"
  </Step>

  <Step title="Enable/Disable OCR">
    Toggle OCR based on your needs:

    * Enabled = Full text extraction (5 tokens/page)
    * Disabled = Process as-is (1 token/page)
  </Step>

  <Step title="Confirm Cost">
    Review page count and total cost before proceeding
  </Step>

  <Step title="Upload">
    Click upload to begin processing
  </Step>
</Steps>

### Matter Upload with OCR

When uploading documents to a Matter:

<Steps>
  <Step title="Open Matter Wizard or Add Documents">
    Start the document addition process
  </Step>

  <Step title="Drop Files">
    Each file is analyzed individually
  </Step>

  <Step title="Per-File OCR Settings">
    Configure OCR for each file:

    * Auto-enabled for likely scanned documents
    * Can override per file
  </Step>

  <Step title="Review Total Cost">
    See combined cost for all documents
  </Step>

  <Step title="Upload All">
    Process all documents with configured settings
  </Step>
</Steps>

## OCR Processing Options

### Processing Modes

<Tabs>
  <Tab title="Anchor-Only (Default)">
    * Extracts text and identifies section structure
    * Faster processing
    * Lower cost
    * Good for standard document analysis
  </Tab>

  <Tab title="Full Processing">
    * Complete text extraction with formatting
    * Preserves tables and layouts
    * Extracts attestations (signatures, stamps)
    * Higher fidelity, longer processing
  </Tab>
</Tabs>

### Attestation Extraction

OCR in Full Processing mode identifies and extracts:

* **Signatures** - Handwritten signatures with location
* **Stamps** - Official stamps, seals, notary marks
* **Certifications** - Notary certificates, apostilles
* **Initials** - Page initials and annotations

These appear as special "attestation cards" in the document view.

<Info>
  Attestations are displayed in distinctive amber-themed cards with clear visual indicators showing what was detected and where.
</Info>

<Frame>
  <img src="https://mintcdn.com/bononetwork/iesyQdBb07VQFOI9/images/attestation.png?fit=max&auto=format&n=iesyQdBb07VQFOI9&q=85&s=dfc727915dc0179517df8eea9dc39a0f" alt="OCR Attestation card showing extracted signature details" width="1009" height="225" data-path="images/attestation.png" />
</Frame>

<p className="text-sm text-muted-foreground mt-2">
  *An attestation card showing a detected signature with description of the signatory and signature style.*
</p>

## Understanding OCR Results

### Text Quality

OCR quality depends on:

| Factor           | Impact                                 |
| ---------------- | -------------------------------------- |
| **Scan quality** | Higher DPI = better results            |
| **Document age** | Older/faded documents harder to read   |
| **Font clarity** | Standard fonts easier than handwriting |
| **Contrast**     | Good black/white contrast helps        |
| **Skew**         | Straight pages process better          |

### Review Recommendations

After OCR processing:

1. **Spot-check critical sections** - Verify important text extracted correctly
2. **Check numbers and dates** - These are often OCR weak points
3. **Review parties and names** - Ensure proper nouns are correct
4. **Validate tables** - Complex tables may need manual review

<Warning>
  Always review OCR'd documents for accuracy before relying on analysis results. OCR is highly accurate but not perfect.
</Warning>

## Cost Considerations

### Standard vs. OCR Processing

| Processing          | Cost     | Best For                             |
| ------------------- | -------- | ------------------------------------ |
| **Fast (Standard)** | Free     | Digital documents with embedded text |
| **OCR**             | 5 tokens | Scanned documents, images            |

OCR processing costs 5 tokens per document, regardless of page count. This covers the AI-powered text extraction and cleanup.

### Pre-Upload Estimation

Arbiter shows costs before you upload:

* Whether the document requires OCR
* Total estimated cost
* Requires explicit confirmation before processing

## Troubleshooting

<AccordionGroup>
  <Accordion title="OCR text is garbled or incorrect">
    * Check original scan quality
    * Try uploading a higher-resolution scan
    * Some handwritten text may not OCR well
    * Consider manual correction for critical sections
  </Accordion>

  <Accordion title="OCR not detecting all text">
    * Very faint text may not extract
    * Multi-column layouts can be challenging
    * Unusual fonts may have issues
    * Ensure document isn't encrypted/protected
  </Accordion>

  <Accordion title="Document uploaded without OCR when needed">
    * Delete the document
    * Re-upload with OCR enabled
    * You cannot retroactively add OCR
  </Accordion>

  <Accordion title="OCR taking very long">
    * Large documents (100+ pages) take time
    * Complex layouts slow processing
    * High-resolution images take longer
    * Typical: 1-2 minutes per 10 pages
  </Accordion>

  <Accordion title="Attestations not appearing">
    * Ensure "Full Processing" mode was used
    * Attestations must be clearly visible in scan
    * Very stylized signatures may not detect
  </Accordion>
</AccordionGroup>

## Best Practices

<Card title="Scan at High Resolution" icon="arrows-maximize">
  300 DPI minimum for text documents. Higher resolution = better OCR accuracy.
</Card>

<Card title="Straighten Before Scanning" icon="align-center">
  Ensure documents are not skewed. Many scanners have auto-straightening features.
</Card>

<Card title="Use Contrast" icon="circle-half-stroke">
  Good black text on white background works best. Avoid colored papers when possible.
</Card>

<Card title="Verify Critical Content" icon="check-double">
  Always manually verify numbers, dates, and party names after OCR processing.
</Card>

## Supported File Types

### For OCR Processing

| Type       | Extensions                     | Notes                       |
| ---------- | ------------------------------ | --------------------------- |
| **PDF**    | .pdf                           | Scanned/image-based PDFs    |
| **Images** | .png, .jpg, .jpeg, .gif, .webp | Photos of documents         |
| **TIFF**   | .tif, .tiff                    | Common for multi-page scans |

### Standard Processing (No OCR Needed)

| Type          | Extensions             |
| ------------- | ---------------------- |
| **PDF**       | .pdf (with text layer) |
| **Word**      | .docx, .doc            |
| **Text**      | .txt                   |
| **Rich Text** | .rtf                   |

## Next Steps

<CardGroup cols={2}>
  <Card title="Document Analysis" icon="file-magnifying-glass" href="/guides/document-analysis">
    Analyze your OCR'd documents
  </Card>

  <Card title="Matter Mode" icon="briefcase" href="/guides/matter-mode">
    Add scanned documents to matters
  </Card>
</CardGroup>
