TIL / A headless-LibreOffice fallback for documents an AI service can't read

A headless-LibreOffice fallback for documents an AI service can't read

Document AIAzureBackend

The problem

A document-intelligence service (in this case, Azure AI Document Intelligence) is tuned for PDFs. Some DOCX files it can’t read directly, even though the file itself opens fine in a word processor. Failing the whole upload on that is a bad experience for a file that’s perfectly readable.

The fix

On a DOCX parse failure, convert the file to PDF with headless LibreOffice and retry the same document-intelligence call against the converted file instead of the original.

libreoffice --headless --convert-to pdf --outdir /tmp/converted input.docx
try:
    result = await doc_intelligence.analyze(docx_bytes)
except UnsupportedFormatError:
    pdf_bytes = convert_to_pdf(docx_bytes)  # shells out to the command above
    result = await doc_intelligence.analyze(pdf_bytes)

Gotcha

Headless LibreOffice is not thread-safe against a shared user profile - running two conversions concurrently against the same profile directory corrupts both. Give each conversion its own isolated profile directory (-env:UserInstallation=file:///tmp/<job-id>) rather than serializing every conversion through a single lock.