Mistral OCR / Document AI - PDF Chunking

OCR

Some documents can be very large and exceed the size (MB) and page limits (count) of what models and services like Mistral OCR and Document AI can provide in a single inference call.

When encountering such restrictions, we recommend chunking your documents. We provide a full end-to-end script here, below we will dig into how it works and some examples to get you started splitting large documents into smaller chunks, and then calling the model endpoint via our SDK.

This can be easily modified to work with existing document processing pipelines, and integrated into many of our cookbooks.

Why chunk at all?#

Mistral OCR and all our Document-AI features enforce two limits:

LimitReasonTypical value
Max file sizeRequests take server memory & upload bandwidth~50 MB
Max page countLong documents cost compute & risk timeouts~1000 pages

Scanned books at high resolution can easily reach these limits. When encountering such restrictions, the fix is to split the document into smaller chunks that each fit, process each chunk separately, then stitch the results back together.

Setup#

For this cookbook, we need two libraries:

  • PyPDF2 — read, count pages, and split PDFs.
  • mistralai - our Mistral AI SDK

Our provided end-to-end script leverages httpx and tenacity with a raw REST implementation to provide a baseline for similar implementations across different languages and services.

!pip install -q PyPDF2 mistralai

Edit the value below. Get a Mistral API key from the Mistral console.

from mistralai.client import Mistral

api_key = "API_KEY"
client = Mistral(api_key)

Encoding a PDF into Base64#

We provide three methods of file uploading for the OCR and Document AI services:

  • Cloud Upload: Organisations can upload their files to our cloud service; once uploaded, a secured URL is returned that lets you perform OCR.
  • Public URL: Provide any publicly accessible URL, and OCR is performed on the document directly.
  • Base64: Encode a document or image in base64 format and send it to the service.

In this cookbook, we will be leveraging Base64, allowing you to OCR any document without any extra cloud-upload steps. To learn about the other approaches, visit our OCR Documentation.

import base64

def encode_pdf_to_base64(pdf_path: str) -> str:
    """Encode the contents of a PDF file to a base64 string.

    Reads the raw bytes of the file and converts them to an ASCII-safe
    base64 string that can be embedded in a JSON request body.
    """
    try:
        # Open in binary mode ('rb') — PDFs are binary, not text.
        with open(pdf_path, "rb") as pdf_file:
            pdf_data = pdf_file.read()
            return base64.b64encode(pdf_data).decode("utf-8")
    except FileNotFoundError:
        print(f"File not found: {pdf_path}")
        raise
    except IOError as e:
        print(f"Error reading file {pdf_path}: {e}")
        raise
    except Exception as e:
        print(f"Error encoding PDF to base64: {e}")
        raise

Chunking#

To split the document, we need a way to count the pages. Using PyPDF, we can make a function like count_pdf_pages:

from PyPDF2 import PdfReader, PdfWriter
from PyPDF2.errors import PdfReadError

def count_pdf_pages(pdf_path: str) -> int:
    """Count the number of pages in a PDF file.

    Uses PyPDF2's PdfReader, which parses the PDF structure without
    rendering each page — so it's fast even for large documents.
    """
    try:
        with open(pdf_path, "rb") as pdf_file:
            pdf_reader = PdfReader(pdf_file)
            return len(pdf_reader.pages)
    except FileNotFoundError:
        print(f"File not found: {pdf_path}")
        raise
    except PdfReadError as e:
        print(f"Error reading PDF structure: {e}")
        raise
    except Exception as e:
        print(f"Error counting PDF pages: {e}")
        raise

Due to the file size limits, we also require a way of estimating the memory size of each document, we will make a function called get_pdf_size_in_mb returning the size in mb.

import os

def get_pdf_size_in_mb(pdf_path: str) -> float:
    """Get the size of a PDF file in MB.

    os.path.getsize() returns bytes; we divide by 1024*1024 to convert
    to megabytes for a human-friendly comparison against MAX_SIZE_MB.
    """
    try:
        size_in_bytes = os.path.getsize(pdf_path)
        return size_in_bytes / (1024 * 1024)  # Convert to MB
    except FileNotFoundError:
        print(f"File not found: {pdf_path}")
        raise
    except OSError as e:
        print(f"Error getting file size: {e}")
        raise

The next step is the heart of the chunking logic. It walks through the PDF in steps of max_pages and creates a new PdfWriter for each chunk.

def split_pdf_by_pages(pdf_path: str, max_pages: int) -> list:
    """Split a PDF into multiple PdfWriters with a maximum number of pages.

    Returns a list of PdfWriter objects (in-memory).
    """
    try:
        with open(pdf_path, "rb") as pdf_file:
            pdf_reader = PdfReader(pdf_file)
            pdf_writers = []

            # Step through pages in blocks of `max_pages`.
            for i in range(0, len(pdf_reader.pages), max_pages):
                pdf_writer = PdfWriter()
                # Add every page in the current block to this writer.
                for page_num in range(i, min(i + max_pages, len(pdf_reader.pages))):
                    pdf_writer.add_page(pdf_reader.pages[page_num])
                pdf_writers.append(pdf_writer)

        return pdf_writers
    except FileNotFoundError:
        print(f"File not found: {pdf_path}")
        raise
    except PdfReadError as e:
        print(f"Error reading PDF structure: {e}")
        raise
    except Exception as e:
        print(f"Error splitting PDF: {e}")
        raise

Why split by pages and not by bytes? Pages are the natural unit of OCR - the API returns results per page. Splitting by bytes would risk cutting a page in half, corrupting the document. Size is only a secondary guard.

Send to the OCR / Document-AI API#

This function leverages the mistralai SDK and implements retries so transient errors (such as rate limits) are handled automatically with exponential backoffs.

import time

def send_to_ocr(base64_pdf: str, retries: int = 10):
    """Send the base64 encoded PDF to Mistral OCR.

    Retries with exponential backoff.
    """
    for i in range(retries):
        try:
            ocr_response = client.ocr.process(
                model="mistral-ocr-latest",
                document={
                    "type": "document_url",
                    "document_url": f"data:application/pdf;base64,{base64_pdf}"
                },
                table_format="html", # default is None
                # extract_header=True, # default is False
                # extract_footer=True, # default is False
                include_image_base64=True
            )
            return ocr_response
        except Exception as e:
            time.sleep(min(4*2**i,32))
    print(e)
    raise

Process the PDF — the orchestrator#

Next, we need to create our orchestrator. For this, we first need to create a function that, given a PDF and our two limits, outputs a list of PDF chunks, all respecting the constraints.

import math

def split_pdf_recursive(pdf_path: str, max_size_mb: float, max_pages: int,
                         output_dir: str = ".", prefix: str = "chunk",
                         _depth: int = 0) -> list:
    """Recursively split a PDF into on-disk chunks that fit both limits.

    Returns a list of file paths to the leaf chunks (each fitting both the
    size and page limits, or a single page that cannot be split further).
    """
    indent = "  " * _depth
    size_mb = get_pdf_size_in_mb(pdf_path)
    num_pages = count_pdf_pages(pdf_path)
    print(f"{indent}PDF: {os.path.basename(pdf_path)} | size: {size_mb:.2f} MB | pages: {num_pages}")

    # Base case: fits both limits — no split needed.
    if size_mb <= max_size_mb and num_pages <= max_pages:
        print(f"{indent}→ fits limits, no split needed")
        return [pdf_path]

    # Edge case: a single page already exceeds the size limit — cannot split further.
    if num_pages <= 1:
        print(f"{indent}⚠ single page exceeds size limit ({size_mb:.2f} MB); left as-is")
        return [pdf_path]

    # Recursive case: choose a chunk size that respects both constraints.
    chunks_for_pages = math.ceil(num_pages / max_pages)
    chunks_for_size = math.ceil(size_mb / max_size_mb)
    num_chunks = max(chunks_for_pages, chunks_for_size)
    pages_per_chunk = math.ceil(num_pages / num_chunks)

    print(f"{indent}→ exceeds limits; splitting into {num_chunks} chunk(s) of ~{pages_per_chunk} page(s)")
    pdf_writers = split_pdf_by_pages(pdf_path, pages_per_chunk)

    chunk_paths = []
    for i, pdf_writer in enumerate(pdf_writers):
        split_pdf_path = os.path.join(output_dir, f"{prefix}_{_depth}_{i}.pdf")
        try:
            with open(split_pdf_path, "wb") as split_pdf_file:
                pdf_writer.write(split_pdf_file)
            # Recurse: re-measures this chunk and splits again if still too large.
            leaf_paths = split_pdf_recursive(
                split_pdf_path, max_size_mb, max_pages,
                output_dir=output_dir, prefix=prefix, _depth=_depth + 1
            )
            chunk_paths.extend(leaf_paths)
        except Exception as e:
            print(f"{indent}Error splitting part {i + 1}: {e}")
            # Clean up a partially-written chunk file on failure.
            try:
                if os.path.exists(split_pdf_path):
                    os.remove(split_pdf_path)
            except OSError:
                pass

    return chunk_paths

We are almost ready to perform or OCR, for this example, we will pick one of the longest research papers available on Arxiv with over 10k pages: Finite-Dimensional Lie Algebras and Their Representations for Unified Model Building

#@title File Download
import urllib.request, os
url = "https://arxiv.org/pdf/1511.08771"
urllib.request.urlretrieve(url, "1511.08771.pdf")
print("Saved:", os.path.abspath("1511.08771.pdf"))

This cell ties everything together:

from tqdm import tqdm

PDF_FILE = "1511.08771.pdf" # One of the longest research papers on arxiv
MAX_SIZE_MB = 25 # API limit is 50
MAX_PAGES = 500 # API limit is 1k

responses = []  # Collect every API response here for later inspection

# Recursively split the PDF into chunks that all fit both limits, then OCR each.
chunk_paths = split_pdf_recursive(PDF_FILE, MAX_SIZE_MB, MAX_PAGES)

print(f"Split into {len(chunk_paths)} chunk(s). Sending each to OCR...")

for i, chunk_path in tqdm(enumerate(chunk_paths)):
    try:
        base64_pdf = encode_pdf_to_base64(chunk_path)
        response = send_to_ocr(base64_pdf)
        responses.append(response)
        print(f"Part {i + 1} processed successfully")
    except Exception as e:
        # One chunk failing shouldn't kill the whole run — log and continue.
        print(f"Error processing part {i + 1}: {e}")
    finally:
        # Clean up the chunk file (unless it's the original, un-split PDF).
        if chunk_path != PDF_FILE:
            try:
                if os.path.exists(chunk_path):
                    os.remove(chunk_path)
            except OSError as e:
                print(f"Could not delete temporary file {chunk_path}: {e}")

print(f"Done. {len(responses)} response(s) collected in `responses`.")

Why try/finally for cleanup? The finally block runs whether the try succeeded or raised. This guarantees the temp file is deleted even if the API call crashes - preventing disk fill-up over many runs.

Done, below we can take a look at the responses.

import json

for i, response in enumerate(responses):
    print(f"--- Response {i + 1} ---")
    try:
        # .json() parses the response body into a Python dict.
        print(json.dumps(response.json(), indent=2)[:2000])
    except Exception:
        # If the body isn't valid JSON, fall back to raw text.
        print(response.text[:2000])