ImgIngImgIng · A DataDance product中文EN日本語한국어DEESPTFROpen PDF content extraction
ImgIng / Extract PDF content
TEXT + IMAGE OCR

Extract native PDF text and every word inside mixed images

Native text is read directly, while every eligible page image is recognized with PP-OCRv6. Overlap with hidden OCR layers is deduplicated.

Native text + OCRMixed text and imagesReading-order reconstructionHTML / TXT / Markdown

Schnelle Fakten zur Fähigkeit

Diese Fakten beschreiben das aktuelle Produkt und nicht die noch nicht ausgelieferte Roadmap-Arbeit.

Text sources
Native PDF text plus image OCR
Modelle
Auto / Fast / Professional / Ultimate
Mixed layout
Merges text and images by page coordinates
Ausgänge
Long HTML / TXT / Markdown

Vom DataDance-Produkt- und Technikteam geprüft · Veröffentlicht · Aktualisiert

Beenden Sie den Vorgang in drei Schritten

Bestätigen Sie wichtige Einstellungen vor dem Export; der Verarbeitungsort wird stets mitgeteilt.

01

Wählen Sie ein PDF

Read text coordinates, image positions and shared resources per page.

02

Choose an OCR tier

Auto uses Professional on desktop and lower-memory Fast on mobile; desktop can also select Fast, Professional or Ultimate.

03

Review and export

Merge native text, images and non-duplicate OCR into a self-contained long page.

Does it work with text PDFs and scans?

Yes. Native words come from PDF text operators; scans and images inside mixed pages use ImgIng PP-OCRv6. Both streams are merged by page coordinates.

How are OCR models selected and loaded?

Auto chooses Professional on desktop and Fast on mobile; desktop users can choose Fast, Professional or Ultimate. On first use the selected model is downloaded and verified from multiple model CDN sources, automatically falls back when one fails, and is then reused from the browser cache.

What if a scan already has an invisible OCR layer?

Normalized OCR output is compared with the page’s native text. Existing phrases are not inserted twice, while unmatched text and the source image remain available for review.

Why export long-page HTML?

HTML keeps page markers, paragraphs, images and recognized image text together. TXT and Markdown provide lighter content-only alternatives.

Updated 2026-08-30 · Live capability detection inside the tool is authoritative

Häufig gestellte Fragen

Diese sichtbaren Antworten entsprechen dem aktuellen Produktverhalten und strukturierten Daten.

Is text inside images recognized?

Yes. Every eligible page image is OCR-processed and merged with native text.

Will hidden OCR text duplicate?

Overlapping normalized text is removed automatically.

Does processing upload the PDF?

No. PDF parsing, OCR, layout reconstruction and export stay local.

Entdecken Sie ImgIng

Auf jeder Aufgabenseite werden echte Einstellungen, Einschränkungen und Formathinweise dokumentiert – keine durch Schlüsselwörter ausgetauschten Duplikate.