Update, September 2026The blocker in this thread was the second one: docx4j had no page layout model, so it couldn't know which row of a table lands on which page. It has one now, so here is how the "Table ABC (continued)" split looks today.
If the output is a docxFrom 17.1.1,
org.docx4j.model.pagination.Paginate lays the document out with Apache FOP (docx4j-export-fo on the classpath; since 17.1.0 that layout follows Word's rules, including for tables: row keeps, repeating header rows, rows that can't split) and reports which page every paragraph starts on - and that includes the paragraphs inside table cells:
- Code: Select all
PaginationMap map = Paginate.compute(wordMLPackage, null);
// for a row: page of its first cell's first paragraph
int page = map.getPageIndex(paraIdOf(firstParagraphIn(tr)));
So the algorithm is:
- paginate;
- walk the table's rows; the first row whose page is later than the table's start page is your split point;
- cut the w:tbl there: a second w:tbl with a copy of w:tblPr and w:tblGrid, the remaining rows moved into it, and your "Table ABC (continued)" paragraph in between (plus a copy of the header row if you want that too);
- repeat for the second table, since it may itself span pages.
Two things to be careful of, one of them from my 2013 answer. Vertically merged cells: don't split inside a merge group - if the row at the split point has a cell with w:vMerge (continue), move the split up to the row that started the merge, or the second table begins with a dangling continuation. And the split itself changes the layout (you've added a paragraph, and the second table starts fresh), so paginate again afterwards and check nothing moved; in practice it settles in one extra pass, but the check is cheap.
If the output is a PDFYou may not need to split at all. Mark the header row as a repeating header (w:tblHeader, "Repeat as header row" in Word); docx4j's PDF output honours it on every page the table continues onto. What that doesn't give you is the literal "(continued)" text - that would need XSL-FO's retrieve-table-marker, which docx4j doesn't emit - so if the wording matters, use the docx-side split above and then convert.
Caveat as always: the layout is FOP's, using the fonts FOP can see, so page boundaries are close to Word's rather than identical. For a table split that is usually fine, since Word re-flows whatever you give it and the split row is a sensible place either way.