WooCommerce products are rarely organized by one flat label. A product can belong to a hierarchy of categories, have one or more brands, and expose attributes such as color, size, material, or compatibility. A Vespa index should preserve that structure instead of reducing every taxonomy value to a single string.
This matters for search, filtering, navigation, autocomplete, and analytics. If taxonomy data is modeled correctly, an agency can support queries such as “black hiking boots from Acme,” category landing pages, brand pages, and faceted navigation without maintaining a separate search-specific taxonomy system.
Identify the WooCommerce taxonomies before indexing
WooCommerce stores product categories and tags as WordPress taxonomies. Categories are hierarchical, while tags are normally flat. Product attributes are represented through taxonomies as well, commonly with names such as pa_color and pa_size. Brands are not part of WooCommerce core, so the brand taxonomy depends on the plugin or custom implementation.
Common brand taxonomy slugs include product_brand, brand, and pwb-brand. An integration should discover the actual taxonomy configuration rather than assuming one slug. The WordPress REST API can expose registered taxonomies through endpoints such as /wp-json/wp/v2/taxonomies, while products are available through /wp-json/wc/v3/products when authenticated WooCommerce API access is configured.
For each product, collect at least:
- The taxonomy term ID and display name for every brand.
- The product category IDs, names, and ancestor relationships.
- Attribute taxonomy names, term IDs, and display values.
- The taxonomy slug or a stable internal identifier.
- The product’s primary category, if the store uses one for breadcrumbs or SEO.
Do not use a display name as the only identifier. Two terms can have the same name, and a term can be renamed without changing its identity. Store IDs for filtering and names for presentation.
Model taxonomy data in the Vespa document
A practical product document separates exact filter fields from searchable text fields. Numeric IDs and normalized slugs are useful for deterministic filtering. Names can be indexed for matching and returned in summaries.
schema product {
document product {
field product_id type int {
indexing: attribute | summary
}
field title type string {
indexing: index | summary
}
field brand_ids type array<int> {
indexing: attribute | summary
}
field brand_names type array<string> {
indexing: index | summary
}
field category_ids type array<int> {
indexing: attribute | summary
}
field category_names type array<string> {
indexing: index | summary
}
field category_path_ids type array<string> {
indexing: attribute | summary
}
field attribute_values type array<string> {
indexing: index | summary
}
}
}
The exact field definitions should reflect the application’s query patterns and Vespa version. For high-volume filtering, numeric IDs in attribute fields are generally preferable to filtering on human-readable names. Names can still be indexed for free-text queries and returned to the frontend.
Preserve category hierarchy
A category such as “Running Shoes” may be nested below “Shoes” and “Sports.” Sending only the leaf term makes it difficult to retrieve all descendants when a shopper selects an ancestor category. During ingestion, expand each product’s category assignments to include its ancestors.
For example, a product assigned to:
Sports > Shoes > Running Shoes
can receive the following category ID values:
category_ids: [10, 24, 87]
category_path_ids: ["10", "10/24", "10/24/87"]
The expanded category_ids field allows a category page for “Shoes” to filter products with category ID 24, including products assigned only to descendants. The path field is useful when the application needs to distinguish branches containing terms with the same name.
Use stable separators and normalized values for path identifiers. Do not build paths from translated names if the path will be used as a filter. Names can change, contain punctuation, or differ between language versions.
Represent brands separately from categories
Brands are typically a many-to-many relationship: a product can have one brand, several brands, or no brand. Store brand IDs as an array even if the current store allows only one brand. That avoids a schema migration if the catalog rules change later.
Keep the brand label in a separate field so the search result can render it without another WordPress request:
brand_ids: [301]
brand_names: ["Acme Outdoor"]
If the catalog supports localized brand names, use locale-specific fields or a structured document model. Avoid indexing translated labels into one unqualified field when a query must be restricted to a particular language.
Index product attributes consistently
WooCommerce attributes may be global taxonomies or product-specific custom attributes. Global attributes usually have reusable term IDs, while custom attributes may exist only inside one product. Normalize both sources into a consistent representation.
For free-text matching, a flattened field can contain values such as:
attribute_values: ["color black", "size large", "material leather"]
For precise faceting, use dedicated fields when the attribute is important to the storefront:
field color_ids type array<int> {
indexing: attribute | summary
}
field color_names type array<string> {
indexing: summary
}
field size_values type array<string> {
indexing: attribute | summary
}
Do not rely on an analyzed text field for exact attribute filtering. For example, “Stainless Steel” may be tokenized into multiple terms, and matching behavior can vary from the intended equality test. Store a normalized identifier or value for filtering and a readable label for display.
Ingest taxonomy changes as document updates
A product document should be rebuilt when its category, brand, or attribute assignments change. The integration can use WooCommerce webhooks for product updates and supplement them with scheduled reconciliation. A full taxonomy synchronization is also needed when a term is renamed, deleted, merged, or moved in the hierarchy.
When a category or brand changes, there are two common strategies:
- Re-fetch and update every affected product document.
- Maintain a separate taxonomy lookup and resolve labels at query or presentation time.
The first strategy is simpler and keeps search results self-contained. The second can reduce repeated product writes in very large catalogs, but it requires additional consistency handling. Most WooCommerce agencies should begin with denormalized product documents and add a taxonomy change queue if catalog size requires it.
Use a deterministic Vespa document ID, such as:
id:product:12345
Make ingestion idempotent. Replaying the same WooCommerce event should produce the same taxonomy arrays and should not create duplicate documents.
Filter products by brand and category
Once IDs are stored as attributes, a search request can combine full-text matching with taxonomy filters. The exact request syntax depends on the query API used by the application, but the logical filter is equivalent to:
where title contains "jacket"
and brand_ids contains 301
and category_ids contains 24
In a Vespa YQL request, the application can express the same constraints using query parameters or a generated query tree. Keep user-entered text separate from trusted taxonomy values. Brand and category IDs should be validated against the store’s catalog before they are inserted into a query.
For multiple selected brands or categories, decide whether the selections mean “any” or “all.” A product matching any selected brand generally uses an OR expression. Requiring a product to contain every selected brand uses multiple conditions and is uncommon for ordinary brand navigation.
Generate facets from the indexed fields
Facet counts should come from the same filtered result set used for the product list. Numeric taxonomy IDs are useful for grouping, while the application can map IDs to labels from the indexed summary fields or a taxonomy cache.
grouping=all(
group(brand_ids)
each(output(count()))
)
For category navigation, grouping on expanded category IDs returns counts for ancestor categories as well as directly assigned leaf categories. If the user interface needs the complete tree, maintain the taxonomy hierarchy separately and merge the returned counts into that tree.
Be careful with facet counts on multi-value fields. A product with several brands contributes to each matching brand bucket. This is usually the desired behavior, but it should be documented when reporting catalog totals or conversion metrics.
Test taxonomy indexing with realistic catalog cases
Before production rollout, test products with no brand, multiple brands, multiple category branches, duplicate term names, deleted terms, translated labels, and custom attributes. Verify that:
- A parent category includes products assigned to its descendants.
- A brand filter does not match a similarly named brand.
- Renaming a term changes the displayed label without changing the filter identifier.
- Removing a category from WooCommerce removes it from the Vespa document.
- Facet counts agree with the active text and taxonomy filters.
- Products with missing or malformed taxonomy data remain searchable.
Log the WooCommerce product ID, taxonomy IDs, ingestion timestamp, and synchronization event ID. These values make it possible to trace a wrong facet or missing product from the storefront back to the source payload.