Chunking should begin with an understanding of the elements in a document, not with an arbitrary character limit. For a WooCommerce agency, this distinction matters because product guides, setup manuals, refund policies, API references, and support exports all contain different structures. A heading, table, warning, code sample, and paragraph should not be treated as interchangeable text.
What an element represents
An element is a meaningful unit produced during parsing. Depending on the parser and output format, an element may represent a title, heading, paragraph, list item, table, image description, code block, page header, page footer, or other document component.
Before creating chunks with LlamaParse and Undefined, inspect the parsed output and identify which fields describe:
- Content: the visible text or structured data.
- Type: whether the item is a heading, paragraph, table, list, code block, or another element.
- Order: the position of the element in the original document.
- Hierarchy: the relationship between a heading and the content that follows it.
- Location: page numbers, coordinates, or source sections where available.
- Metadata: document identifiers, URLs, product names, language, and other retrieval fields.
Why element inspection matters in WooCommerce
A support article might contain a heading such as Configure shipping zones, several explanatory paragraphs, a numbered procedure, and a table of shipping methods. If these elements are split without regard to their structure, a retrieved chunk may contain a shipping-rate row without its column headings or a procedure step without the conditions that make it valid.
The same problem appears in product documentation. A chunk containing “Enable the setting” is not useful unless it retains the heading, the preceding explanation, and the relevant product or extension name. Element inspection provides the context needed to keep those relationships intact.
Inspect the parsed representation first
Do not assume that the parser has produced the structure you expect. Save a small parsed document and inspect its element types, text, ordering, and metadata before designing chunk rules.
for index, element in enumerate(parsed_document.elements):
print({
"index": index,
"type": element.type,
"text": element.text[:120],
"metadata": element.metadata,
})
The exact property names depend on the LlamaParse output format and the Undefined integration you use. The important practice is to inspect the actual result rather than writing chunking logic against assumptions. Confirm whether tables are returned as complete elements, whether list items are separate, and whether headings contain their level.
Classify elements by their chunking behavior
A useful first pass is to classify elements into groups with different handling rules.
Structural elements
Headings establish scope. A heading such as Refunds for digital products should normally be attached as context to the paragraphs and lists beneath it. Preserve the heading hierarchy when possible, including the parent heading for deeply nested sections.
prose elements
Paragraphs are usually safe to combine when they belong to the same section and do not exceed the target size. Preserve paragraph boundaries so that sentence and formatting context is not lost.
Atomic elements
Tables, code blocks, warnings, formulas, and image captions often need to remain intact. Splitting them by character count can remove headers, syntax, or qualification statements. If an atomic element is too large, split it using a structure-aware rule and retain an explicit continuation marker.
List elements
Keep a list with its introductory sentence when that sentence defines the meaning of the list. For a procedure, preserve step order and avoid combining unrelated steps from different subsections.
Build a normalized element model
Different parser outputs can use different names for the same concepts. Normalizing them before chunking makes the rest of the pipeline easier to test and maintain.
normalized = {
"index": index,
"kind": "heading", # heading, paragraph, list, table, code, image, other
"text": text,
"heading_level": level,
"page": page_number,
"section_path": section_path,
"source_id": document_id,
}
Store the normalized model separately from the original parser response. The original response remains useful for debugging, while the normalized representation gives the chunker a stable input. Avoid silently discarding fields such as page numbers or table structure; they can be valuable when an agency needs to show the source of an answer.
Use element boundaries to form chunks
A practical chunking process can follow these steps:
- Start a new section context when a heading is encountered.
- Attach the current section path to subsequent elements.
- Add compatible paragraphs and list items until the target size is reached.
- Keep tables, code blocks, and warnings intact whenever possible.
- Split only oversized elements, using their internal structure rather than raw character positions.
- Store the source location and element types with every resulting chunk.
For example, a chunk for a WooCommerce shipping guide might contain the section path Settings > Shipping > Shipping zones, two explanatory paragraphs, and a three-step procedure. A separate chunk might contain the shipping-method table, with its caption and column headings preserved.
Choose chunk size after understanding structure
Token or character limits are constraints, not the primary definition of a chunk. First determine which elements belong together; then use the size limit to decide where a structurally safe split is required. This avoids a common failure mode in which every chunk fits the limit but loses the context needed for retrieval or answer generation.
For agency documentation, different content types may justify different limits. Short configuration instructions can remain compact, while a table describing dozens of shipping methods may require a dedicated chunking strategy. Keep the limit configurable and record the applied strategy in chunk metadata.
Validate element assumptions with test documents
Use a small test set that includes the document patterns your agency processes regularly:
- A product guide with nested headings.
- A PDF containing a multi-page table.
- A refund policy with numbered clauses.
- An API guide containing code blocks.
- A scanned document requiring OCR.
- A support export containing repeated headers and footers.
Review the parsed elements and resulting chunks manually. Check that headings remain connected to their content, table headers are not separated from values, repeated page furniture is removed or marked, and OCR errors do not create misleading section boundaries.
Metadata to retain for retrieval
Each chunk should carry enough metadata to support filtering, debugging, and citations. Common fields include the source document, URL, page number, section path, product or extension name, language, element types, and a stable chunk identifier. If a chunk was created by splitting an oversized table or code block, record that fact as well.
This metadata allows a retrieval system to distinguish, for example, a chunk about WooCommerce Subscriptions from a similarly worded chunk about a payment gateway. It also gives the agency a reliable way to trace an answer back to the source document.
A pre-chunking checklist
- Have the parser’s actual element types been inspected?
- Are headings associated with the correct section content?
- Are tables, code, warnings, and lists handled explicitly?
- Are headers, footers, navigation, and duplicated OCR text identified?
- Are source locations and section metadata preserved?
- Are oversized elements split using structure-aware rules?
- Have representative WooCommerce documents been tested end to end?