Skip to content

Text with background missing from output #275

Closed
@rsaite

Description

@rsaite

Referencing #225 which was incorrectly marked as fixed even though the bug still persists.

PDF for reproduction: https://bmjopensem.bmj.com/content/bmjosem/1/1/e000050.full.pdf

Consider the box on the middle right of the first page with blue and grey background ("Strenghts and limitations of this study * Free weight resistance training [..]"). That text is embedded in the pdf and can be extracted, e.g. using PyMuPDF. However, when doing a simple markdown conversion with PyMuPDF4LLM, that text is missing from the output.

#225 was closed because the bug was supposedly fixed with release 0.0.18. However, the bug can still be observed in release 0.0.18 and also in the current release 0.0.24.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions