Document Extractor — PDF to text, Markdown, tables
Turns PDF URLs into one dataset row per page: plain text, Markdown, tables as row arrays, and an honest has_text_layer flag for scanned pages. Built to sit behind any crawler that downloads PDFs but does not read them.
50%
users