TeleOCR: A 1.2B Vision-Language Model for Structured Document Parsing
TeleOCR is an open-source, approximately 1.2-billion-parameter vision-language model for parsing bot 2026-10-3 09:33:11 Author: hackernoon.com(查看原文) 阅读量:3 收藏

TeleOCR is an open-source, approximately 1.2-billion-parameter vision-language model for parsing both born-digital documents and camera-captured pages. It is maintained by XingChen-AGI, uses the Transformers library, and is loaded through AutoProcessor and AutoModel. Its central distinction is that it targets geometric distortion and structured content—especially tables and formulas—as well as ordinary text extraction. The README describes deformation-aware document modeling, adaptive sampling, and content-structure decoupled learning; it does not specify the model’s image-resolution limit, training dataset size, training steps, VRAM requirement, or inference speed. Treat the benchmark results as evidence for document-parsing workloads, not as a guarantee for every language, scan quality, or production pipeline.

The project reports strong results on several document benchmarks, including an overall score of 96.87 on OmniDocBench v1.6 and 88.53 on Wild_OmniDocBench. It also reports first place in the ICDAR 2026 Sci-ImageMiner Challenge. These are project-reported evaluations; the supplied material does not provide enough detail to independently assess all benchmark protocols or reproduce the scores.

Best use cases

Parsing phone photos of pages with perspective or curvature distortion. The model is designed to handle camera-captured documents in the same framework as digital documents. Its geometry-aware modeling and deformation-aware learning target the layout errors that can cascade through pipelines dependent on accurate layout detection. The README says it parses distorted pages from DocUNet and DIR300 without dewarping preprocessing or a dedicated rectification model.

Extracting text, tables, and reading order from mixed-layout documents. OmniDocBench v1.6 reports 96.87 overall, 97.05 Table TEDS, and 98.52 Table TEDS-S for the model. Those results make it a candidate for documents where preserving table structure matters alongside text extraction. The reported read-order edit score is 0.122; lower is better, and several alternatives score lower on that metric.

Parsing formulas and tables as structured content. The model’s content-structure decoupled learning explicitly targets table structures and formula grammars. This is relevant when downstream systems need more than a plain text transcript. Benchmark results are not uniformly best across every formula metric, however, so validate on the notation and document types in your own corpus.

Processing degraded documents. On PureDocBench, the reported overall scores are 77.47 for digital-degraded documents and 70.85 for real-degraded documents. The model leads the listed comparison rows on those overall scores, but the benchmark table also shows that other models can lead individual submetrics. Use a representative test set before replacing a specialized OCR pipeline.

Running a compact, specialized document parser rather than a large general VLM. At about 1.2B parameters, it is much smaller than the 30B–241B general VLMs listed in the OmniDocBench comparison. That makes it a more focused option to evaluate for document parsing, though the README gives no measured latency, memory footprint, or hardware baseline to establish deployment cost.

Limitations

The supplied documentation does not state a maximum input resolution, supported image dimensions, maximum page count, context length, batch-size guidance, or inference speed. It also gives no VRAM requirement or hardware benchmark. The 1.2B parameter count alone is not enough to calculate actual inference memory: image preprocessing, precision, framework behavior, and generation length also affect it.

The README presents benchmark scores but does not describe detailed failure cases, language coverage, handwriting performance, confidence calibration, or robustness by document subtype. It also does not provide a complete reproducibility package in the supplied text: training dataset names and sizes, training steps, compute budget, and evaluation settings are absent. Do not infer those details from the reported scores.

The Dr.DocBench table is mixed rather than a clean win across all metrics. TeleOCR has the highest listed overall score (67.96), but its formula CDM is 0.02, below the listed alternatives’ 0.04 and 0.21, and its Table TEDS (64.97) is below MinerU 2.5 Pro’s 67.75. On OmniDocBench v1.6, it leads overall and table scores among the listed specialized VLMs, but OvisOCR2 has better text edit, formula CDM, and read-order edit scores. Metric direction matters: the tables mark text edit and read-order edit as lower-is-better, while overall, formula CDM, and table scores are higher-is-better.

The license is Apache-2.0, which is generally permissive, but the supplied material does not include the full license text or separate terms for any hosted service. Check the repository’s license and your organization’s compliance requirements before deployment. The README does not state safety policies, known biases, or fine-tuning instructions.

How it compares

TeleOCR. This is the same model family and description under the StarDoc-AI listing, so the benchmark and technical claims here apply to that listing rather than representing a distinct competitor. Choose it when you want the documented 1.2B document-parsing model and its reported strength on distorted pages, tables, and mixed digital/camera inputs; the supplied information does not establish a separate speed, cost, or quality difference between the listings.

Unlimited-OCR. The provided comparison material identifies this as Baidu’s “Unlimited OCR Works” and describes a focus on one-shot, long-horizon parsing, but gives no parameter count, benchmark scores, hardware needs, or implementation details. Choose TeleOCR when you need the specific reported results and documented approach for camera distortion and structured tables; consider Unlimited-OCR if its long-horizon parsing workflow matches your task, then compare both on your own documents because the supplied information does not support a quantitative quality, speed, or cost ranking.

Qianfan-OCR. Qianfan-OCR is described as a 4B-parameter end-to-end document intelligence model that combines parsing, layout analysis, and document understanding in one vision-language architecture. TeleOCR is the smaller option at about 1.2B parameters and has reported benchmark strength on distorted documents and tables; Qianfan-OCR may fit workflows that need broader document understanding, but the provided excerpt does not include comparable benchmark or latency figures to establish which is more accurate or faster.

GLM-OCR. GLM-OCR is described as a multimodal OCR model built on a GLM-V encoder-decoder architecture, with a CogViT visual encoder, token downsampling, a GLM-0.5B language decoder, Multi-Token Prediction loss, and full-task reinforcement learning. In the supplied OmniDocBench v1.6 table, TeleOCR scores higher overall (96.87 vs. 95.22) and on table metrics (97.05 TEDS vs. 92.83; 98.52 TEDS-S vs. 95.39), while GLM-OCR has a higher formula CDM score (97.18 vs. 96.36). Pick based on whether overall/table parsing or formula performance is more important; no speed or cost comparison is provided.

PaddleOCR-VL-1.6. This is a 0.9B specialized model in the supplied OmniDocBench table. TeleOCR scores higher overall (96.87 vs. 96.33) and on table TEDS and TEDS-S (97.05/98.52 vs. 94.76/97.11), while PaddleOCR-VL-1.6 has a higher formula CDM score (97.49 vs. 96.36) and a lower text edit score (0.033 vs. 0.027 is actually worse, since lower is better; TeleOCR leads there). Choose TeleOCR for the reported overall and table results or its stated camera-distortion focus; choose PaddleOCR-VL-1.6 if its formula score is more relevant to your evaluation. The parameter counts suggest a smaller model for PaddleOCR-VL-1.6, but no measured speed or memory comparison is supplied.

Qianfan-OCR and GLM-OCR are the only other alternatives in the supplied internal-link list with enough description or benchmark data for a direct technical comparison. The remaining linked alternatives lack comparable details in the provided excerpts, so a quality, speed, or cost ranking would be unsupported.

Technical specifications

TeleOCR is a specialized vision-language model for document parsing, with approximately 1.2B parameters. The README describes a unified approach to digital and camera-captured documents. Its stated methods include Multi-node Consensus Voting (MCV) for automatic pseudo-label generation, geometry-aware document modeling for camera-captured documents, Curvature-Guided Douglas-Peucker Sampling (CGDP), image-to-image self-verification for automatic data refinement, a progressive four-stage training pipeline, and content-structure decoupled learning for tables and formulas. The abstract characterizes the approach as deformation-aware learning, adaptive sampling for complex layouts, and explicit modeling of formula grammars and table structures.

Confirmed implementation and deployment details:

  • Task and pipeline: image-to-text / image-text-to-text document parsing.
  • Model size: approximately 1.2B parameters.
  • Library: Transformers; the README’s installation command also installs PyTorch and Pillow.
  • Loading API: AutoProcessor and AutoModel.
  • Input shown in the example: a PIL image.
  • Output representation shown in the example: the code defines content blocks with a type, bounding box, optional angle, and optional content; it also defines table cells with text, row/column offsets, and spans.
  • Table representation: the example includes OTSL tokens , , , , , and for parsing table structure.
  • License: Apache-2.0.
  • Reported deployment/community notes: the README mentions a community GGUF conversion with llama.cpp support and a community deployment report on Ascend 910B. These are community contributions, not stated as the official Transformers inference path.
  • Not specified: maximum image resolution, accepted file formats beyond the PIL-image example, supported precision or quantization in the official release, VRAM, latency, batch size, training dataset and size, training steps, and compute budget.

The README’s benchmark tables report the following key results:

  • OmniDocBench v1.6: overall 96.87; text edit 0.027; formula CDM 96.36; table TEDS 97.05; table TEDS-S 98.52; read-order edit 0.122.
  • Wild_OmniDocBench: overall 88.53; text edit 0.1173; formula CDM 88.26; table TEDS 89.05; table TEDS-S 92.14; read-order edit 0.2011.
  • PureDocBench: clean overall 86.90; digital-degraded overall 77.47; real-degraded overall 70.85. The abstract separately reports 78.41 overall on PureDocBench, so the supplied material contains different PureDocBench figures; check the evaluation protocol before comparing them.
  • Dr.DocBench Challenge: overall 67.96; text edit 0.1903; formula CDM 0.02; table TEDS 64.97; order edit 0.398.
  • ICDAR 2026 Sci-ImageMiner: first place in the listed table, with RMS 17.23, TEDS 66.39, and weighted score 41.81.

Model inputs and outputs

Inputs

  • Image: the example loads an image with Pillow and passes it to the processor. The README does not specify supported image file formats, color modes, dimensions, or resolution limits.
  • Document content: the intended inputs include digital documents and camera-captured documents, including geometrically distorted pages.
  • Batching and multi-page input: not specified in the supplied documentation.

Outputs

  • Structured parsed content: the example defines content blocks with type, bbox, optional angle, and optional content.
  • Table structure: the example defines cells with text, start/end row and column offsets, and row/column spans. It also includes OTSL structural tokens for table parsing.
  • Post-processing: the README’s example imports html, itertools, json, and re and begins helper functions for extracting OTSL tokens and text. This indicates that table output may need parsing into application-specific structures; the supplied excerpt does not show the complete inference and post-processing code or the exact serialized model output.

Getting started

Install the dependencies listed in the README:

pip install transformers torch pillow

The README’s example uses AutoProcessor and AutoModel and a PIL image. The supplied excerpt does not include the model-loading identifier, processor call, generation method, or complete output-decoding code, so a fully executable inference snippet cannot be reconstructed without inventing details. Use the model repository’s current usage example to confirm the exact checkpoint identifier and inference call before integrating it.

Frequently asked questions

Q: Can I use TeleOCR commercially?

A: The model is listed under the Apache-2.0 license. Review the license file and any applicable repository or service terms before commercial deployment.

Q: What hardware or VRAM do I need to run TeleOCR?

A: The supplied README does not state a VRAM requirement or minimum hardware. It reports approximately 1.2B parameters and mentions community deployment on Ascend 910B, but that does not establish a minimum configuration.

Q: How does TeleOCR compare with GLM-OCR for table parsing?

A: On the supplied OmniDocBench v1.6 table, TeleOCR scores 97.05 Table TEDS and 98.52 Table TEDS-S, compared with GLM-OCR’s 92.83 and 95.39. GLM-OCR has the higher formula CDM score, so evaluate both if formula recognition is central.

Q: What are the known failure modes or quality issues?

A: The supplied material does not document specific failure cases. Its benchmark results vary by metric: for example, it does not lead every formula or read-order metric, and the Dr.DocBench table shows a low formula CDM score of 0.02.

Q: Can I fine-tune TeleOCR, and what framework does it use?

A: The README identifies Transformers as the library and shows loading through AutoProcessor and AutoModel, but it does not provide fine-tuning instructions, training code, or a supported fine-tuning recipe.

Q: What input format does TeleOCR expect?

A: The example passes a Pillow image to the processor. The README does not specify accepted file formats, image dimensions, or resolution limits.

Q: How fast is inference, and what batch sizes are practical?

A: The supplied documentation gives no inference-time, throughput, or batch-size measurements. Benchmark the model on your target hardware and document sizes.

Q: Is TeleOCR still actively maintained?

A: The README lists releases and project updates in August and September 2026, including a rename from NaviDC-OCR to TeleOCR and a statement that subsequent iterations will use the TeleOCR name. The supplied information does not establish a maintenance schedule beyond those announcements.

Research paper: NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents is the related research-paper page provided for this model. The README identifies the technical report as arXiv:2608.12898.

This is a simplified guide to an AI model called TeleOCR maintained by XingChen-AGI.


文章来源: https://hackernoon.com/teleocr-a-12b-vision-language-model-for-structured-document-parsing?source=rss
如有侵权请联系:admin#unsafe.sh