Extract text as Markdown, inferring headings from font size. Works best on text-based PDFs, not scans.