Skip to content

fix: prevent unparse crash on split surrogate entities - #36

Open
sabbirexist wants to merge 1 commit into
rjriajul:devfrom
sabbirexist:fix/unparse-split-surrogate-entities
Open

sabbirexist wants to merge 1 commit into
rjriajul:devfrom
sabbirexist:fix/unparse-split-surrogate-entities

Conversation

@sabbirexist

Copy link
Copy Markdown
Contributor
  • Fixes a UnicodeDecodeError in HTML and Markdown unparse when an entity boundary falls inside a UTF-16 surrogate pair.
  • Adds clamp_to_code_point() to safely widen misaligned boundaries to the complete Unicode code point.
  • Keeps all valid/aligned entity behavior unchanged.
  • Adds 13 regression tests covering the affected cases.
  • Full test suite: 14862 passed, 35 skipped.

@rjriajul

rjriajul commented Oct 8, 2026

Copy link
Copy Markdown
Owner

Thanks for this. The half-emoji crash is real, and the new tests catch it (7 of them fail on dev and pass on this branch).

One case I hit while testing: when two entities overlap after the widening, the Markdown output puts the tags in the wrong order. Examples:

  • a😀b with bold at offset 1 length 1 and italic at offset 2 length 1 gives a**__😀**__b
  • 😀x with italic at offset 0 length 1 and bold at offset 1 length 1 gives __**😀__**x

HTML comes out fine for both. Could you add a test for overlapping entities in Markdown and make sure the output stays well formed? Let me know when it's pushed and I'll take another look.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants