Skip to content

[Bug]: WordParser drops soft line breaks and tabs inside DOCX paragraphs #2839

Description

@Xscaperrr

Prerequisites

  • I have searched the existing issues and discussions, and this is not a duplicate.
  • This is a bug, not a usage question. (For questions, please use Discussions instead.)

Background / Description

WordParser.parse() silently removes soft line breaks and tabs inside a
nonempty DOCX paragraph. The surrounding text is concatenated without a
separator, corrupting the content passed to RAG chunking and indexing.

A soft line break, such as one inserted with Shift+Enter in Word, stays
inside the same paragraph. It is represented by a w:br element; a tab is
represented by w:tab. The reproduction below creates both using python-docx
and confirms that they survive saving and reopening the DOCX.

Expected: preserve the internal separators:

'Name\tAlice\nDepartment\tEngineering'

Actual: the fields and lines are joined together:

'NameAliceDepartmentEngineering'

This concerns separators inside a single paragraph, rather than blank lines
between separate paragraphs.

Error Messages

Expected: 'Name\tAlice\nDepartment\tEngineering'
Actual:   'NameAliceDepartmentEngineering'
AssertionError


The assertion is from the reproduction. The parser itself returns the altered
text without raising an error.

Steps to Reproduce

From a checkout of AgentScope, install the project and the Word parser dependency:

python -m pip install -e . python-docx

Save the following as repro.py and run python repro.py. It creates a real
DOCX in memory and requires no API key or external service.

import asyncio
import io

from docx import Document
from agentscope.rag import WordParser


async def main():
    expected = "Name\tAlice\nDepartment\tEngineering"
    document = Document()
    document.add_paragraph(expected)
    buffer = io.BytesIO()
    document.save(buffer)
    data = buffer.getvalue()

    # Check that the DOCX itself retains the tabs and soft line break.
    assert Document(io.BytesIO(data)).paragraphs[0].text == expected

    sections = await WordParser(include_image=False).parse(data, "example.docx")
    assert len(sections) == 1
    actual = sections[0].content.text
    print(f"Expected: {expected!r}")
    print(f"Actual:   {actual!r}")
    assert actual == expected


asyncio.run(main())

Environment

  • AgentScope: 2.0.8, editable install at 083cbd1975c7b5ed055c69b0d19c150727e7f606.
  • Python: 3.11
  • python-docx: 1.2.0.
  • OS: Windows.
  • On that checkout, the existing WordParser tests pass:
    python -m pytest tests/rag_parser_test.py -k WordParser -q → 14 passed.
    The reproduction above fails.
  • The same reproduction also fails when loading _word.py from PR fix(rag): keep text inside nested Word tables #2759's
    final commit (4a16631de82fce3e3d15fa6cc059422ccc6b4107) and upstream main
    (a38821287f35e9e45ed193d9d864cb46f263c946, checked on 2026-09-25) against
    the local environment. These additional checks exercised the parser file,
    not the full test suite of those revisions.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

BugSomething doesn't worktriage/confirmedVerified: the reported defect exists

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions