Skip to content

[Bug]: WordParser drops the content of block-level w:sdt content controls (automatic TOC, controlled tables) #2843

Description

@sxh313

Prerequisites

  • I have searched the existing issues and discussions, and this is not a duplicate.
  • This is a bug, not a usage question. (For questions, please use Discussions instead.)

Background / Description

WordParser.parse walks the document body and only recognises two kinds of
child (rag/_parser/_word.py:297-332, on main = a3882128):

for element in doc.element.body:          # :297
    if isinstance(element, CT_P):         # :298
        ...
    elif isinstance(element, CT_Tbl):     # :332
        ...

Anything else that is a direct child of w:body is skipped without a word.
One of those things is w:sdt — a block-level content control, which is
how Word stores its automatic Table of Contents, its "Bibliography", every
Building Blocks entry, and any developer-mode block control. Its content
lives one level down:

<w:sdt>
  <w:sdtContent>
    <w:p>…</w:p>      <!-- paragraphs and tables, invisible to the loop -->
    <w:tbl>…</w:tbl>
  </w:sdtContent>
</w:sdt>

So a document whose table of contents sits in the body parses as if that
region were blank — the whole TOC (or a table of controlled content) never
reaches the index.

The asymmetry is the tell: an inline w:sdt (one nested inside a w:p)
is read correctly today, because python-docx hands the enclosing paragraph to
_extract_text_from_paragraph, whose para._element text walk reaches the
control's runs. Only the container form is lost, and that is the form Word
uses for its automatic TOC.

Reproduction

No network, no API key, no Word installation — the documents are built in
memory with python-docx, which is already this parser's dependency.

# repro_word_sdt.py
import asyncio
import io

from docx import Document as DocxDocument
from docx.oxml import parse_xml

from agentscope.rag import WordParser

_W = 'xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"'

TOC_SDT = f"""
<w:sdt {_W}>
  <w:sdtPr><w:id w:val="123456"/></w:sdtPr>
  <w:sdtContent>
    <w:p><w:r><w:t>Table of Contents</w:t></w:r></w:p>
    <w:p><w:r><w:t>1  Introduction ....... 1</w:t></w:r></w:p>
    <w:p><w:r><w:t>2  Results ............ 4</w:t></w:r></w:p>
  </w:sdtContent>
</w:sdt>
"""

SDT_WITH_TABLE = f"""
<w:sdt {_W}>
  <w:sdtContent>
    <w:p><w:r><w:t>Controlled heading</w:t></w:r></w:p>
    <w:tbl>
      <w:tr>
        <w:tc><w:p><w:r><w:t>cell-a</w:t></w:r></w:p></w:tc>
        <w:tc><w:p><w:r><w:t>cell-b</w:t></w:r></w:p></w:tc>
      </w:tr>
    </w:tbl>
  </w:sdtContent>
</w:sdt>
"""

NESTED_SDT = f"""
<w:sdt {_W}>
  <w:sdtContent>
    <w:p><w:r><w:t>outer text</w:t></w:r></w:p>
    <w:sdt>
      <w:sdtContent>
        <w:p><w:r><w:t>inner text</w:t></w:r></w:p>
      </w:sdtContent>
    </w:sdt>
  </w:sdtContent>
</w:sdt>
"""

INLINE_SDT = f"""
<w:p {_W}>
  <w:r><w:t>Before control </w:t></w:r>
  <w:sdt>
    <w:sdtContent>
      <w:r><w:t>controlled value</w:t></w:r>
    </w:sdtContent>
  </w:sdt>
</w:p>
"""

CASES = [
    ("Word automatic TOC", TOC_SDT, ["Table of Contents", "Introduction"]),
    (
        "content control holding a table",
        SDT_WITH_TABLE,
        ["Controlled heading", "cell-a"],
    ),
    ("nested content control", NESTED_SDT, ["outer text", "inner text"]),
]


def _bytes_of(doc: DocxDocument) -> bytes:
    buf = io.BytesIO()
    doc.save(buf)
    return buf.getvalue()


def _build(container_xml: str) -> bytes:
    doc = DocxDocument()
    doc.add_paragraph("Paragraph before.")
    after = doc.add_paragraph("Paragraph after.")
    # Put the container between the two paragraphs: ``w:sectPr`` has to
    # stay the last body child.
    after._element.addprevious(parse_xml(container_xml))
    return _bytes_of(doc)


async def _text_of(data: bytes) -> str:
    sections = await WordParser(include_image=False).parse(data, "r.docx")
    return "\n".join(
        s.content.text for s in sections if hasattr(s.content, "text")
    )


async def main() -> None:
    total_missing = 0
    for label, xml, expected in CASES:
        data = _build(xml)
        text = await _text_of(data)
        gone = [fragment for fragment in expected if fragment not in text]
        total_missing += len(gone)
        print(f"--- {label} ---")
        print(
            "  paragraphs python-docx exposes for this doc:",
            len(DocxDocument(io.BytesIO(data)).paragraphs),
        )
        print("  parser text:", repr(text))
        print("  lost inside the container:", gone if gone else "nothing")

    print("--- inline content control, for contrast ---")
    doc = DocxDocument()
    doc.add_paragraph("Paragraph before.")
    holder = doc.add_paragraph("")
    holder._element.addprevious(parse_xml(INLINE_SDT))
    print("  parser text:", repr(await _text_of(_bytes_of(doc))))

    print(f"\ntotal block-level fragments lost: {total_missing}")


asyncio.run(main())

Output on main (a3882128), Python 3.11 — all three container forms lose
both of their fragments:

--- Word automatic TOC ---
  paragraphs python-docx exposes for this doc: 2
  parser text: 'Paragraph before.\nParagraph after.'
  lost inside the container: ['Table of Contents', 'Introduction']
--- content control holding a table ---
  paragraphs python-docx exposes for this doc: 2
  parser text: 'Paragraph before.\nParagraph after.'
  lost inside the container: ['Controlled heading', 'cell-a']
--- nested content control ---
  paragraphs python-docx exposes for this doc: 2
  parser text: 'Paragraph before.\nParagraph after.'
  lost inside the container: ['outer text', 'inner text']
--- inline content control, for contrast ---
  parser text: 'Paragraph before.\nBefore control controlled value'

total block-level fragments lost: 6

Note the last block: the inline control is read fine (Before control controlled value), which is what makes the three above a container-walking
gap rather than a general blindness to w:sdt.

Nothing here is a parsing subtlety — the control is sitting in plain sight as
a direct child of the body. Dumping the body tags for each case:

--- Word automatic TOC ---
  direct body children: ['p', 'sdt', 'p', 'sectPr']
  w:sdt present as a body child: True
--- content control holding a table ---
  direct body children: ['p', 'sdt', 'p', 'sectPr']
  w:sdt present as a body child: True
--- nested content control ---
  direct body children: ['p', 'sdt', 'p', 'sectPr']
  w:sdt present as a body child: True

Expected behaviour

A block-level content control is a container, not content: the loop should
descend into its w:sdtContent and treat the children it finds there exactly
as if they had been body children in that position — recursively, since a
control can hold another control. Then all three cases above print their
fragments and the total is 0, while the existing handling of w:p,
w:tbl, images, separate_table and table_format is untouched.

Section boundaries matter here, so "descend" must keep the surrounding
buffering behaviour: a controlled paragraph between two ordinary paragraphs
should merge into the same Section the way its neighbours do, and a
controlled table should still obey separate_table.

Not covered by this issue

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    triage/confirmedVerified: the reported defect exists

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions