Why does a PDF paste with a break on every line?
A PDF describes where each line of text sits on the page, not where paragraphs begin and end. When you copy, the viewer joins the fragments it finds and ends each visual line with a newline. The result is text that wraps at the width of the original page, with the wrong line lengths in your email, note or document.
Two problems get mixed together here, and they need different fixes: the hidden characters that come along (soft hyphens, odd spaces), and the hard-wrapped lines. The first is what the cleaner above handles. The second is a find-and-replace job.
Step 1: remove the hidden characters
Paste the text into the cleaner. It removes soft hyphens and zero-width characters, converts non-breaking spaces to ordinary spaces, and standardizes line endings. Do this first so the patterns below do not trip over invisible characters. More detail is on the PDF text cleaner page.
Step 2: join the lines
| Goal | Find (regex) | Replace with |
|---|---|---|
| Repair words split by a hyphen at line end | (\w)-\n(?=[a-z]) | $1 |
| Join lines inside a paragraph | (?<!\n)\n(?!\n) | One space |
In JavaScript, run the first two in order: text.replace(/(\w)-\n(?=[a-z])/g, "$1").replace(/(?<!\n)\n(?!\n)/g, " "). Editors that support regular expressions (VS Code, Sublime Text, Notepad++ in regex mode) accept the same patterns.
Two honest limits. The hyphen pattern will also join a genuine hyphenated word that happens to break at the hyphen, such as well-known, so skim the result. And if the PDF has no blank line between paragraphs, no pattern can tell where one paragraph ends, so you will need to mark paragraph breaks by hand first.
Step 3: do it in Word or Google Docs
- Word: open Find and Replace (Ctrl+H). Replacing ^p with a space also joins paragraphs, so first replace ^p^p with a placeholder such as @@, then replace ^p with a space, then swap the placeholder back to ^p^p.
- Google Docs: Edit, Find and replace, tick "Match using regular expressions" for the hyphen pattern. Lookbehind patterns may not be supported there, so use the placeholder approach from the Word bullet to join lines.
What if the text still looks wrong?
Columns interleaved, page numbers in the middle of sentences, and scanned pages are layout and OCR problems. Copy one column at a time, use your PDF reader's export or Reading Mode where it exists, or run OCR on scans first. The wider background is in why text pastes weird and the general spacing fix.