An Empirical Study of VLM Pipelines for Long-Document QA
arXiv:2609.29933v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts. Deploying them means choosing how to feed the document to the model, which retriever…