markitdown.ai

markitdown.ai - A Markdown converter for documents and web pages that turns PDF, Office files, images and HTML into AI-ready Markdown

Launched today

markitdown.ai is a conversion tool that transforms PDFs, DOCX, PPTX, XLSX, images, and public web URLs into clean, structurally accurate Markdown. It is built for AI developers, data teams, and knowledge management professionals who need predictable, low-noise text output for RAG pipelines, LLM workflows, and agent systems without manual cleanup.

3ViewsAI DataFreemiumLow-CodeDocument ProcessingNLPLarge Language ModelRAG

Core Capabilities

markitdown.ai supports conversion from 7 input formats and 3 output formats, all running on a shared server-side parsing engine that produces consistent results regardless of client device. Supported input formats:

  • PDF: Native text PDFs parse directly, multi-column reading order is rebuilt, scanned PDFs go through automatic OCR

  • DOCX / legacy DOC: Preserves headings, lists, tables, and comments from Word files

  • PPTX / legacy PPT: Converts slide decks page by slide, retains titles as headings and bullet points

  • XLSX / legacy XLS: Each workbook sheet becomes a Markdown table under its sheet name

  • PNG / JPG: Screenshots and scanned images processed via OCR, optional AI image understanding for signed-in users

  • HTML: Strips scripts and styles, preserves article structure, links and tables from saved web pages

  • URL: Fetches public web pages server-side, blocks private and internal addresses to avoid leaking restricted content

Supported reverse output formats, which run fully locally in the browser without uploading user text:

  • Markdown to HTML: Generates semantic, standard HTML with support for GitHub tables, task lists, and code highlighting

  • Markdown to plain text: Flattens structure into readable unformatted text for chat windows, support tickets, and form inputs

  • Markdown to PDF: Renders a local print preview that users can save using their browser's native print function

Workflow Options

The product offers three distinct ways to run conversions, all connected to the same user account and shared credit balance:

  1. Web app (browser): No sign-in required for anonymous trial use, supports single file uploads or URL pastes, with up to 10 MB file size and 25 page limits on the free trial tier. Users can preview rendered Markdown, copy content, or download .md files directly.

  2. Batch mode: For paid plan users, supports processing multiple files in one run, with per-file status tracking and a saved conversion library.

  3. Developer API: Uses the exact same parsing pipeline as the web interface. It supports synchronous conversion under 60 seconds, asynchronous conversion for longer files with polling or webhook completion notifications, scoped API keys, and full activity logs. The public API reference is available at /docs/api, and the OpenAPI spec can be accessed directly at /v1/openapi.json.

Audience and Use Cases

The tool is built for these core teams and use cases:

  • RAG and AI pipeline teams: Convert source documents before chunking and embedding to eliminate layout noise, preserve heading hierarchy and table structure, and keep retrieval results consistent across thousands of files.

  • Researchers and analysts: Convert academic papers, whitepapers, and slide decks into searchable, editable Markdown to avoid manual retyping of tables and preserve correct multi-column reading order from dense PDFs.

  • Operations teams: Process contracts, invoices, policies, and internal documents without manual copy-paste cleanup. Stored Markdown history simplifies review, comparison, and reuse of recurring document work.

  • Developer automation teams: Embed conversion directly into ingestion jobs, internal tools, and document workflows, with predictable output that feeds reliably into CMS systems, knowledge bases, and support bots.

  • Agent builders: Add Markdown conversion as a tool for AI agents, so they can process supported files and public web URLs without handling low-level parsing or OCR logic.

Pricing and Usage Limits

The anonymous, no-signup trial allows 3 conversions per hour and 10 conversions per day, with a 10 MB maximum file size cap and 25 page maximum per file. Credit-based pricing applies across all plans: every standard text page or OCR processed page costs 1 credit, while opt-in AI image understanding costs 5 credits per image. There is no difference in per-page cost between a text-layer document and a scanned document. Paid subscriptions start at $20 per month, and increase size limits up to 200 MB per file and 200 pages per file on the Pro plan. Paid plans also unlock batch processing, API access, and longer stored history for conversion outputs.

Important Limitations

  • OCR recognition quality is dependent on the source image resolution, so 300 dpi scans produce far more accurate text output than low-resolution phone photos.

  • Private and internal web addresses are blocked for URL-to-Markdown conversion to prevent unauthorized access to non-public content.

  • Structured JSON extraction, audio transcription, and video transcription are listed as roadmap features and are not yet available as shipped functionality.

Comments

Comments

No comments yet. Be the first to share your thoughts!