Listening to a document instead of reading it
There are two quite different reasons to want this. The first is circumstance: a commute, a walk, cooking, anything where your eyes are busy and your attention is not. A long report becomes something you can get through rather than something waiting for a moment you never find.
The second is accessibility, and it matters more. For people with visual impairments, dyslexia or conditions that make sustained reading difficult, having a document read aloud is not a convenience but the difference between the document being available and not.
There is also a proofreading use that surprises people. Hearing your own writing read back catches awkward phrasing, repetition and missing words far more reliably than reading it again, because your eye skips what it expects and your ear does not.
What works well and what does not
Continuous prose reads well: articles, reports, letters, chapters. The voice handles sentence structure sensibly and the result is easy to follow.
Structured content works less well, because a document's visual organisation carries meaning that speech does not. Tables read as a stream of unrelated values. Headings blur into the text that follows. Footnotes interrupt mid-sentence. For anything where layout does the explaining, listening is a poor substitute for looking.
The document also has to contain real text. A scanned PDF has nothing to read aloud, so OCR is a prerequisite rather than an option, and the audio can only ever be as accurate as the recognition was.
Punctuation is read as prosody rather than spoken, so a document with erratic punctuation produces speech with erratic pacing. Text extracted from a PDF is a common offender here, because line breaks and hyphenation from the original layout survive into the text and the voice pauses in places no human reader would. Cleaning the text before listening makes a substantial difference to how easy it is to follow.
Getting a listenable result
Extract the part you actually want first. Whole documents include cover pages, tables of contents, headers repeated on every page and reference lists, all of which are tedious to listen through and easy to remove beforehand.
Adjust the speed to the material rather than leaving it at default. Familiar prose is comfortable well above normal speaking pace; dense or technical text is not, and slowing down for the difficult section is more effective than replaying it.
Speech synthesis has improved enormously, but it still stumbles on unusual proper nouns, abbreviations and technical vocabulary. If a term matters and comes out wrong, it is worth checking the written document rather than assuming you misheard.
Voices, speed and listening to long documents
Browser speech synthesis has improved considerably, and most systems now offer several voices with genuinely different characteristics. It is worth trying more than the default: some voices handle punctuation and sentence rhythm noticeably better, and one that sounds fine for a paragraph can become grating over an hour.
Accent matters for comprehension more than people expect. A voice matching the English variety you are used to will be easier to follow at speed, and if the document contains local place names or terminology, a regional voice will usually pronounce them more sensibly.
For anything long, break the document into sections rather than generating one continuous stream. It makes it far easier to return to where you stopped, to re-listen to a section you missed, and to skip material you do not need — all of which are awkward in a single two-hour recording with no structure.
Be realistic about retention. Listening suits material you want to absorb broadly — an overview, a narrative, a report you need the gist of. It suits close study much less well, because you cannot easily reread a sentence, hold two passages side by side, or scan back for a definition. For anything you will be tested on or have to act precisely upon, listening is a good first pass and a poor only pass.