Beyond OCR… it understands Arabic documents and their structure
Drag to upload
Full language support with excellence in Arabic and obtaining data that is copyable and directly usable
Technology developed by Misraj AI
Why does Baseer surpass traditional OCR?
Baseer transcends conventional OCR technology. It's a Vision Language Model (VLM) purpose-built to intelligently process Arabic documents and transform them into professionally structured, actionable formats.
Baseer reads texts, tables, equations, and images accurately while maintaining the original formatting
Professional processing of Arabic documents

Expert in Arabic Poetry
Baseer recognizes Arabic poetry and formats it correctly, separating the first hemistich from the second hemistich with three stars (***) to preserve the poetic meter and rhythm.

Advanced Smart Text Extraction
Basir reads documents naturally and smoothly, just like a human would, while preserving the original formatting and text order.

Mathematical Equations Support
Extracts mathematical and scientific equations in professional LaTeX format, ensuring high accuracy when handling scientific and technical content.
Professional Table Processing
It converts tables into organized HTML format ready for use, making them easy to display and process in any application.

Image Location Detection
Identifies the positions of images within a document and inserts special <img> tags in the text to distinguish them from the surrounding content.
Intelligent Handling of Special Elements
Watermarks: Recognizes watermarks (Official Copy, Confidential) and wraps them in <watermark> tags.

Upload, Edit… and Save Your Documents Easily in Word & PDF
Baseer provides a complete document processing environment that allows you to upload your Arabic files and edit their content accurately, with the ability to export the final results in Word and PDF formats according to enterprise-ready standards.
Baseer surpasses leading global models in accuracy benchmarks.
We tested 'Baseer' - our specialized model trained to understand Arabic documents and images and transform them into structured, usable text - and compared it against the world's most powerful models. Baseer demonstrated exceptional superiority in text comprehension and maintaining table structure, delivering the most accurate digital outputs achievable by any AI model today. At Misraj Lab, we take pride in developing 'Baseer' as the premier vision-language model
| Model | WER↓ lower is better | CER↓ lower is better | BLEU↑ higher is better | CHRF↑ higher is better | TEDS↑ higher is better | MARS↑ higher is better |
|---|---|---|---|---|---|---|
| ★ Baseer_v2 | 0.19 | 0.12 | 76.31 | 88.94 | 67 | 77.8 |
| gemini_2.5_pro | 0.37 | 0.31 | 77.92 | 89.55 | 52 | 70.775 |
| Azure Document Intelligence | 0.44 | 0.27 | 62.04 | 82.49 | 42 | 62.245 |
| Dots_ocr | 0.5 | 0.4 | 58.16 | 78.41 | 40 | 59.205 |
| Nanonets_OCR2_3B | 0.78 | 0.71 | 44.29 | 68.39 | 49 | 58.695 |
| GPT-5 | 0.86 | 0.62 | 40.67 | 61.6 | 48 | 54.8 |
| Qwen2_5_vl_32b | 0.76 | 0.59 | 37.62 | 62.64 | 41 | 51.82 |
| Qwen3_VL_8B_Instruct | 0.87 | 0.78 | 32.95 | 54.54 | 49 | 51.77 |
| DeepSeek_OCR | 0.88 | 0.81 | 41.57 | 62.33 | 30 | 46.165 |
| MISTRAL | 0.49 | 0.42 | 52.44 | 71.81 | 17 | 44.405 |
| Qari | 0.76 | 0.64 | 38.59 | 64.5 | 21 | 42.75 |
| Gemma3_12B | 0.96 | 0.8 | 19.75 | 44.53 | 33 | 38.765 |
Key Features
Comprehensive Feature Set
Data Privacy and Security
Your files are processed within a secure enterprise environment, without the need for external models or services

Easy Integration (API / Model)
Seamlessly integrate with your company's systems to automate Arabic document processing and convert them into ready-to-use data, enhancing operational efficiency and reducing costs
Wide Document Support
Baseer, not trained on handwritten documents, effectively handles a range of Arabic documents, from books, magazines, and academic articles to invoices and tabular data...


The First Model for Arabic Text Extraction
The best performing model among all open-source and closed-source models
Direct Document to Markdown Conversion
Preserves formatting, tables, headings, and lists, ready for processing or archiving
Arabic OCR from Baseer — High Accuracy
Baseer provides accurate OCR results in Arabic
Text from a book
PDF - 2 pages
Extracted Text


How Baseer serves different sectors in processing Arabic documents with accuracy and reliability.
Examples of Use Cases
Companies and Organizations
Companies deal daily with large volumes of invoices, reports, and various forms. Baseer enables high-accuracy reading of these Arabic documents, automatically extracts key data, and converts them into ready-to-use information within business systems.
Benefits:
- Accelerate operations and reduce time wasted on manual entry
- Lower operational costs and reduce human errors
- Seamless integration with existing systems such as ERP and accounting systems

Government Entities
Government entities suffer from massive accumulation of paper archives and old documents. Baseer helps convert these files into organized digital archives, while maintaining privacy and enabling on-premise operation without cloud dependency.
Benefits:
- Intelligent archiving of official documents
- Instant search across millions of pages within seconds
- Complete independence from the cloud and full data control

Analysis and Research
Baseer serves researchers and educational institutions in handling scanned or complex content, converting it into structured text that can be easily searched and analyzed.
Benefits:
- Convert scanned notes and lectures into usable digital text
- Archive theses, dissertations, and academic references
- Efficiently analyze large PDF files and extract information from them

Specialized articles to help you maximize the use of text conversion and extraction tools
Insights and articles on text processing and conversion

Start Digitizing Your Arabic Documents Today
A model that outperforms all globally available solutions in processing Arabic content. Simply put, Baseer understands Arabic better than any other model in its field, allowing you to upload your files, edit their content with precision, then save them in Word or PDF formats ready for use within your systems or direct sharing
Choose Your Plan
Choose the plan that fits your needs
FAQs
Frequently Asked Questions
Unlike traditional Optical Character Recognition (OCR) systems that analyze images at the component level and may produce unstructured text with incorrect content order, Baseer relies on an advanced vision-language model that understands the entire image and extracts structured and accurate text while preserving context and layout.