> ## Documentation Index
> Fetch the complete documentation index at: https://docs.lapathoniia.top/llms.txt
> Use this file to discover all available pages before exploring further.

# Rukopys OCR

> Ukrainian handwritten document recognition

<Tabs sync={false}>
  <Tab title="EN">
    ## What is Rukopys?

    Rukopys is a specialised model for recognising **Ukrainian handwritten text**. Supports images and PDF documents.

    **Supported formats:** PNG, JPG, JPEG, WebP, PDF, TIF, TIFF (up to 50 MB)

    **Quota:** 4 images/hour (default). Quota applies **per page** for multi-page files.

    ## Via Chat UI

    Go to [app.lapathoniia.top](https://app.lapathoniia.top), select **Rukopys OCR** model, attach an image or PDF.

    ## Via API

    ```python theme={null}
    import httpx, base64, json

    with open("handwritten.jpg", "rb") as f:
        image_b64 = base64.b64encode(f.read()).decode()

    resp = httpx.post(
        "https://app.lapathoniia.top/chat/stream",
        headers={"Authorization": "Bearer YOUR_API_KEY"},
        json={
            "message": "Recognise text",
            "model": "rukopys-ocr",
            "session_id": "ocr-1",
            "user_id": "user-123",
            "web_search": False,
            "connectors": False,
            "chat_history": [],
            "images": [{"base64": image_b64, "mime_type": "image/jpeg"}],
        },
    )

    result = []
    for line in resp.iter_lines():
        if line.startswith("data: "):
            d = json.loads(line[6:])
            if "text" in d:
                result.append(d["text"])

    regions = json.loads("".join(result))
    ```

    ### Response Format

    ```json theme={null}
    [
      {
        "text": "Dear Sir,",
        "bbox": [120, 45, 380, 70],
        "confidence": 0.94,
        "page": 1
      },
      {
        "text": "We are forwarding the documents",
        "bbox": [120, 80, 520, 105],
        "confidence": 0.91,
        "page": 1
      }
    ]
    ```

    | Field | Description |
    | - | - |
    | `text` | Recognised text of the region |
    | `bbox` | Coordinates \[x1, y1, x2, y2] in pixels |
    | `confidence` | Model confidence (0–1) |
    | `page` | Page number (for PDFs) |

    ## OCR + LLM Pipeline

    Combine Rukopys OCR with a language model — OCR via `/chat/stream`, then extraction with MamayLM.

    ```python theme={null}
    import httpx, base64, json

    def ocr_then_extract(image_path: str, prompt: str, api_key: str) -> dict:
        with open(image_path, "rb") as f:
            b64 = base64.b64encode(f.read()).decode()

        headers = {"Authorization": f"Bearer {api_key}"}

        # Step 1 — OCR
        ocr_resp = httpx.post(
            "https://app.lapathoniia.top/chat/stream",
            headers=headers,
            json={
                "message": "Розпізнай текст",
                "model": "rukopys-ocr",
                "session_id": "ocr-pipeline",
                "user_id": "agent",
                "web_search": False,
                "connectors": False,
                "chat_history": [],
                "images": [{"base64": b64, "mime_type": "image/jpeg"}],
            },
        )
        ocr_text = "".join(
            json.loads(l[6:])["text"]
            for l in ocr_resp.iter_lines()
            if l.startswith("data: ") and "text" in json.loads(l[6:])
        )

        # Step 2 — Extract structured data
        llm_resp = httpx.post(
            "https://app.lapathoniia.top/chat/stream",
            headers=headers,
            json={
                "message": f"{prompt}\n\nText:\n{ocr_text}",
                "model": "MamayLM-Gemma-3-27B-IT-v2.0",
                "session_id": "ocr-pipeline",
                "user_id": "agent",
                "web_search": False,
                "connectors": False,
                "chat_history": [],
            },
        )
        result_text = "".join(
            json.loads(l[6:])["text"]
            for l in llm_resp.iter_lines()
            if l.startswith("data: ") and "text" in json.loads(l[6:])
        )
        return json.loads(result_text)

    # Example: grade extraction from handwritten tests
    result = ocr_then_extract(
        "test.jpg",
        "Extract: student name, date, subject, score (0-100). Return JSON only.",
        api_key="sk-...",
    )
    # → {"student": "Ivan Petrenko", "date": "2024-05-12", "subject": "Math", "score": 87}
    ```

    Other pipeline use cases:

    * `"Extract: date, signatories, amounts. Return JSON only."` — contracts
    * `"Summarize this handwritten note in 2 sentences."` — meeting notes
    * `"Translate to English."` — Ukrainian handwritten → English

    ## Limits

    * **4 images/hour** — default quota per key
    * **50 MB** — maximum file size
    * **50 pages** — maximum pages per PDF/TIFF
    * Check current quota: [`GET /user/limits`](/api/usage-limits)
  </Tab>

  <Tab title="UA">
    ## What is Rukopys?

    Rukopys — спеціалізована модель для розпізнавання **українських рукописних текстів**. Підтримує зображення та PDF.

    **Формати:** PNG, JPG, JPEG, WebP, PDF, TIF, TIFF (до 50 МБ)

    **Квота:** 4 зображення/год. Квота рахується **по сторінках** для багатосторінкових файлів.

    ## Via API

    ```python theme={null}
    import httpx, base64, json

    with open("manuscript.jpg", "rb") as f:
        b64 = base64.b64encode(f.read()).decode()

    resp = httpx.post(
        "https://app.lapathoniia.top/chat/stream",
        headers={"Authorization": "Bearer YOUR_API_KEY"},
        json={
            "message": "Розпізнай текст",
            "model": "rukopys-ocr",
            "session_id": "ocr-1",
            "user_id": "user-123",
            "web_search": False,
            "connectors": False,
            "chat_history": [],
            "images": [{"base64": b64, "mime_type": "image/jpeg"}],
        },
    )
    ```

    ### Response Format

    ```json theme={null}
    [
      {
        "text": "Шановний пане,",
        "bbox": [120, 45, 380, 70],
        "confidence": 0.94,
        "page": 1
      }
    ]
    ```

    ## OCR + LLM Pipeline

    Поєднайте Rukopys OCR з мовною моделлю — OCR через `/chat/stream`, потім витяг через MamayLM.

    ```python theme={null}
    import httpx, base64, json

    def ocr_та_витяг(шлях: str, інструкція: str, api_key: str) -> dict:
        with open(шлях, "rb") as f:
            b64 = base64.b64encode(f.read()).decode()

        headers = {"Authorization": f"Bearer {api_key}"}

        # Крок 1 — OCR
        ocr_resp = httpx.post(
            "https://app.lapathoniia.top/chat/stream",
            headers=headers,
            json={
                "message": "Розпізнай текст",
                "model": "rukopys-ocr",
                "session_id": "ocr-pipeline",
                "user_id": "agent",
                "web_search": False,
                "connectors": False,
                "chat_history": [],
                "images": [{"base64": b64, "mime_type": "image/jpeg"}],
            },
        )
        ocr_text = "".join(
            json.loads(l[6:])["text"]
            for l in ocr_resp.iter_lines()
            if l.startswith("data: ") and "text" in json.loads(l[6:])
        )

        # Крок 2 — Витяг структурованих даних
        llm_resp = httpx.post(
            "https://app.lapathoniia.top/chat/stream",
            headers=headers,
            json={
                "message": f"{інструкція}

    Текст:
    {ocr_text}",
                "model": "MamayLM-Gemma-3-27B-IT-v2.0",
                "session_id": "ocr-pipeline",
                "user_id": "agent",
                "web_search": False,
                "connectors": False,
                "chat_history": [],
            },
        )
        result_text = "".join(
            json.loads(l[6:])["text"]
            for l in llm_resp.iter_lines()
            if l.startswith("data: ") and "text" in json.loads(l[6:])
        )
        return json.loads(result_text)

    # Приклад: витяг оцінок з рукописних контрольних
    result = ocr_та_витяг(
        "test.jpg",
        "Витягни: ім’я учня, дату, предмет, оцінку (0-100). Поверни тільки JSON.",
        api_key="sk-...",
    )
    # → {"student": "Іван Петренко", "date": "2024-05-12", "subject": "Математика", "score": 87}
    ```

    Інші варіанти використання:

    * `"Витягни: дату, підписантів, суми. JSON."` — договори
    * `"Стисло переказ у 2 реченнях."` — нотатки з нарад
    * `"Перекладіть англійською."` — рукопис → англійська
  </Tab>
</Tabs>
