Skip to main content

Fix Formatting

Fix Line Breaks Copied From a PDF

CleanPastedText editorial · Updated

Text workbench

Try it on your text

Cleaning mode

Keeps smart quotes, dashes, and ellipses while cleaning hidden characters.

Try a real example

Each sample contains a problem you cannot see.

Original

Pasted text

0 chars · 0 words

Cleaned

Ready to copy

0 changes

Text stays in this browser

Same words · No AI rewriting · No content logging

Why does a PDF paste with a break on every line?

A PDF describes where each line of text sits on the page, not where paragraphs begin and end. When you copy, the viewer joins the fragments it finds and ends each visual line with a newline. The result is text that wraps at the width of the original page, with the wrong line lengths in your email, note or document.

Two problems get mixed together here, and they need different fixes: the hidden characters that come along (soft hyphens, odd spaces), and the hard-wrapped lines. The first is what the cleaner above handles. The second is a find-and-replace job.

Step 1: remove the hidden characters

Paste the text into the cleaner. It removes soft hyphens and zero-width characters, converts non-breaking spaces to ordinary spaces, and standardizes line endings. Do this first so the patterns below do not trip over invisible characters. More detail is on the PDF text cleaner page.

Step 2: join the lines

GoalFind (regex)Replace with
Repair words split by a hyphen at line end(\w)-\n(?=[a-z])$1
Join lines inside a paragraph(?<!\n)\n(?!\n)One space

In JavaScript, run the first two in order: text.replace(/(\w)-\n(?=[a-z])/g, "$1").replace(/(?<!\n)\n(?!\n)/g, " "). Editors that support regular expressions (VS Code, Sublime Text, Notepad++ in regex mode) accept the same patterns.

Two honest limits. The hyphen pattern will also join a genuine hyphenated word that happens to break at the hyphen, such as well-known, so skim the result. And if the PDF has no blank line between paragraphs, no pattern can tell where one paragraph ends, so you will need to mark paragraph breaks by hand first.

Step 3: do it in Word or Google Docs

  • Word: open Find and Replace (Ctrl+H). Replacing ^p with a space also joins paragraphs, so first replace ^p^p with a placeholder such as @@, then replace ^p with a space, then swap the placeholder back to ^p^p.
  • Google Docs: Edit, Find and replace, tick "Match using regular expressions" for the hyphen pattern. Lookbehind patterns may not be supported there, so use the placeholder approach from the Word bullet to join lines.

What if the text still looks wrong?

Columns interleaved, page numbers in the middle of sentences, and scanned pages are layout and OCR problems. Copy one column at a time, use your PDF reader's export or Reading Mode where it exists, or run OCR on scans first. The wider background is in why text pastes weird and the general spacing fix.

Common questions

Frequently asked questions

Why does text copied from a PDF have a line break on every line?

A PDF is a page-layout format. It places each line of text at fixed coordinates and has no built-in concept of a paragraph. When a viewer copies the text, it usually ends each visual line with a newline, so your pasted text wraps at the width of the printed page instead of the width of your document.

How do I remove line breaks from PDF text but keep paragraphs?

Replace each single newline with a space, and leave a newline that is followed by another newline. In a regex-capable editor or in JavaScript, the pattern (?<!\n)\n(?!\n) matches a lone line break. This only works if your paragraphs are separated by a blank line, which many PDF viewers do not add.

Does CleanPastedText join PDF lines?

No. It removes hidden characters, converts unusual spaces, standardizes line endings and reduces runs of blank lines, but it deliberately keeps ordinary line breaks so lists, addresses and code are not merged by accident. Use the pattern on this page after cleaning if you want lines joined.

How do I fix words split by a hyphen at the end of a line?

First remove any soft hyphens (U+00AD), which the cleaner does. For visible hyphens at a line end, join the two halves only when the second half starts with a lowercase letter, using the pattern (\w)-\n(?=[a-z]). Review the result, because a real hyphenated word such as well-known would also be joined.

Why is there a space or odd character inside PDF words?

Some PDFs place letters as separate positioned fragments, and the viewer guesses where spaces belong. Others use ligature glyphs or non-breaking spaces. A cleaner can normalise the special spaces and hidden characters, but it cannot know the original spelling, so check names, numbers and words with ligatures such as fi.

Continue reading

Related guides & tools