Prerequisites
Background / Description
WordParser.parse walks the document body and only recognises two kinds of
child (rag/_parser/_word.py:297-332, on main = a3882128):
for element in doc.element.body: # :297
if isinstance(element, CT_P): # :298
...
elif isinstance(element, CT_Tbl): # :332
...
Anything else that is a direct child of w:body is skipped without a word.
One of those things is w:sdt — a block-level content control, which is
how Word stores its automatic Table of Contents, its "Bibliography", every
Building Blocks entry, and any developer-mode block control. Its content
lives one level down:
<w:sdt>
<w:sdtContent>
<w:p>…</w:p> <!-- paragraphs and tables, invisible to the loop -->
<w:tbl>…</w:tbl>
</w:sdtContent>
</w:sdt>
So a document whose table of contents sits in the body parses as if that
region were blank — the whole TOC (or a table of controlled content) never
reaches the index.
The asymmetry is the tell: an inline w:sdt (one nested inside a w:p)
is read correctly today, because python-docx hands the enclosing paragraph to
_extract_text_from_paragraph, whose para._element text walk reaches the
control's runs. Only the container form is lost, and that is the form Word
uses for its automatic TOC.
Reproduction
No network, no API key, no Word installation — the documents are built in
memory with python-docx, which is already this parser's dependency.
# repro_word_sdt.py
import asyncio
import io
from docx import Document as DocxDocument
from docx.oxml import parse_xml
from agentscope.rag import WordParser
_W = 'xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"'
TOC_SDT = f"""
<w:sdt {_W}>
<w:sdtPr><w:id w:val="123456"/></w:sdtPr>
<w:sdtContent>
<w:p><w:r><w:t>Table of Contents</w:t></w:r></w:p>
<w:p><w:r><w:t>1 Introduction ....... 1</w:t></w:r></w:p>
<w:p><w:r><w:t>2 Results ............ 4</w:t></w:r></w:p>
</w:sdtContent>
</w:sdt>
"""
SDT_WITH_TABLE = f"""
<w:sdt {_W}>
<w:sdtContent>
<w:p><w:r><w:t>Controlled heading</w:t></w:r></w:p>
<w:tbl>
<w:tr>
<w:tc><w:p><w:r><w:t>cell-a</w:t></w:r></w:p></w:tc>
<w:tc><w:p><w:r><w:t>cell-b</w:t></w:r></w:p></w:tc>
</w:tr>
</w:tbl>
</w:sdtContent>
</w:sdt>
"""
NESTED_SDT = f"""
<w:sdt {_W}>
<w:sdtContent>
<w:p><w:r><w:t>outer text</w:t></w:r></w:p>
<w:sdt>
<w:sdtContent>
<w:p><w:r><w:t>inner text</w:t></w:r></w:p>
</w:sdtContent>
</w:sdt>
</w:sdtContent>
</w:sdt>
"""
INLINE_SDT = f"""
<w:p {_W}>
<w:r><w:t>Before control </w:t></w:r>
<w:sdt>
<w:sdtContent>
<w:r><w:t>controlled value</w:t></w:r>
</w:sdtContent>
</w:sdt>
</w:p>
"""
CASES = [
("Word automatic TOC", TOC_SDT, ["Table of Contents", "Introduction"]),
(
"content control holding a table",
SDT_WITH_TABLE,
["Controlled heading", "cell-a"],
),
("nested content control", NESTED_SDT, ["outer text", "inner text"]),
]
def _bytes_of(doc: DocxDocument) -> bytes:
buf = io.BytesIO()
doc.save(buf)
return buf.getvalue()
def _build(container_xml: str) -> bytes:
doc = DocxDocument()
doc.add_paragraph("Paragraph before.")
after = doc.add_paragraph("Paragraph after.")
# Put the container between the two paragraphs: ``w:sectPr`` has to
# stay the last body child.
after._element.addprevious(parse_xml(container_xml))
return _bytes_of(doc)
async def _text_of(data: bytes) -> str:
sections = await WordParser(include_image=False).parse(data, "r.docx")
return "\n".join(
s.content.text for s in sections if hasattr(s.content, "text")
)
async def main() -> None:
total_missing = 0
for label, xml, expected in CASES:
data = _build(xml)
text = await _text_of(data)
gone = [fragment for fragment in expected if fragment not in text]
total_missing += len(gone)
print(f"--- {label} ---")
print(
" paragraphs python-docx exposes for this doc:",
len(DocxDocument(io.BytesIO(data)).paragraphs),
)
print(" parser text:", repr(text))
print(" lost inside the container:", gone if gone else "nothing")
print("--- inline content control, for contrast ---")
doc = DocxDocument()
doc.add_paragraph("Paragraph before.")
holder = doc.add_paragraph("")
holder._element.addprevious(parse_xml(INLINE_SDT))
print(" parser text:", repr(await _text_of(_bytes_of(doc))))
print(f"\ntotal block-level fragments lost: {total_missing}")
asyncio.run(main())
Output on main (a3882128), Python 3.11 — all three container forms lose
both of their fragments:
--- Word automatic TOC ---
paragraphs python-docx exposes for this doc: 2
parser text: 'Paragraph before.\nParagraph after.'
lost inside the container: ['Table of Contents', 'Introduction']
--- content control holding a table ---
paragraphs python-docx exposes for this doc: 2
parser text: 'Paragraph before.\nParagraph after.'
lost inside the container: ['Controlled heading', 'cell-a']
--- nested content control ---
paragraphs python-docx exposes for this doc: 2
parser text: 'Paragraph before.\nParagraph after.'
lost inside the container: ['outer text', 'inner text']
--- inline content control, for contrast ---
parser text: 'Paragraph before.\nBefore control controlled value'
total block-level fragments lost: 6
Note the last block: the inline control is read fine (Before control controlled value), which is what makes the three above a container-walking
gap rather than a general blindness to w:sdt.
Nothing here is a parsing subtlety — the control is sitting in plain sight as
a direct child of the body. Dumping the body tags for each case:
--- Word automatic TOC ---
direct body children: ['p', 'sdt', 'p', 'sectPr']
w:sdt present as a body child: True
--- content control holding a table ---
direct body children: ['p', 'sdt', 'p', 'sectPr']
w:sdt present as a body child: True
--- nested content control ---
direct body children: ['p', 'sdt', 'p', 'sectPr']
w:sdt present as a body child: True
Expected behaviour
A block-level content control is a container, not content: the loop should
descend into its w:sdtContent and treat the children it finds there exactly
as if they had been body children in that position — recursively, since a
control can hold another control. Then all three cases above print their
fragments and the total is 0, while the existing handling of w:p,
w:tbl, images, separate_table and table_format is untouched.
Section boundaries matter here, so "descend" must keep the surrounding
buffering behaviour: a controlled paragraph between two ordinary paragraphs
should merge into the same Section the way its neighbours do, and a
controlled table should still obey separate_table.
Not covered by this issue
Prerequisites
Background / Description
WordParser.parsewalks the document body and only recognises two kinds ofchild (
rag/_parser/_word.py:297-332, onmain=a3882128):Anything else that is a direct child of
w:bodyis skipped without a word.One of those things is
w:sdt— a block-level content control, which ishow Word stores its automatic Table of Contents, its "Bibliography", every
Building Blocksentry, and any developer-mode block control. Its contentlives one level down:
So a document whose table of contents sits in the body parses as if that
region were blank — the whole TOC (or a table of controlled content) never
reaches the index.
The asymmetry is the tell: an inline
w:sdt(one nested inside aw:p)is read correctly today, because python-docx hands the enclosing paragraph to
_extract_text_from_paragraph, whosepara._elementtext walk reaches thecontrol's runs. Only the container form is lost, and that is the form Word
uses for its automatic TOC.
Reproduction
No network, no API key, no Word installation — the documents are built in
memory with
python-docx, which is already this parser's dependency.Output on
main(a3882128), Python 3.11 — all three container forms loseboth of their fragments:
Note the last block: the inline control is read fine (
Before control controlled value), which is what makes the three above a container-walkinggap rather than a general blindness to
w:sdt.Nothing here is a parsing subtlety — the control is sitting in plain sight as
a direct child of the body. Dumping the body tags for each case:
Expected behaviour
A block-level content control is a container, not content: the loop should
descend into its
w:sdtContentand treat the children it finds there exactlyas if they had been body children in that position — recursively, since a
control can hold another control. Then all three cases above print their
fragments and the total is
0, while the existing handling ofw:p,w:tbl, images,separate_tableandtable_formatis untouched.Section boundaries matter here, so "descend" must keep the surrounding
buffering behaviour: a controlled paragraph between two ordinary paragraphs
should merge into the same
Sectionthe way its neighbours do, and acontrolled table should still obey
separate_table.Not covered by this issue
fix(word): preserve inline tabs and breaks in Word paragraphs) is a different gap in a different function —_extract_text_from_paragraphlosingw:tab/w:br/w:crinside aparagraph. It does not reach block-level
w:sdt, and this issue does nottouch that function.
_extract_table_datacollects inside a cell; the second case above isoutside that path, since the table is a child of the control rather than of
a cell.
w:sdtalready works and needs no change.