WooCommerce categories carry more than display metadata. They provide a stable way to filter products, build landing pages, generate facets, and preserve the structure of a store’s catalog. In Vespa, the most useful design is usually to store category membership directly on product documents and optionally index categories as their own documents when category pages need search, ranking, or independent metadata.
Choose the category representation
There are two common models for WooCommerce category data:
- Embedded category membership: each product contains its category IDs, slugs, names, and ancestor IDs. This is the best model for category filters and product search.
- Category documents: each WooCommerce category becomes a Vespa document. This is useful when category pages need descriptions, images, SEO fields, product counts, or search.
These models can be used together. Product documents answer questions such as “show products in Boots,” while category documents answer questions such as “find categories matching waterproof footwear.”
Store stable IDs and query-friendly values
WooCommerce category names and slugs can change. The numeric WooCommerce category ID is normally the most stable identifier, so use it for product filtering and synchronization. Store slugs and names as additional fields for URLs, display, and text search.
A product document can contain fields like these:
{ "id": "product::581", "fields": { "product_id": 581, "title": "Waterproof Hiking Boot", "category_ids": [18, 42], "category_slugs": ["shoes", "boots"], "category_names": ["Shoes", "Boots"], "category_ancestor_ids": [7, 18, 42] } }
The ancestor field is important for hierarchical navigation. If category 42 is a child of category 18, and category 18 is a child of category 7, products in category 42 can contain all three IDs. A query for category 18 then returns products in that category and its descendants without requiring a recursive query at search time.
Define product fields in Vespa
For exact filtering, category IDs should be attributes. The fast-search setting makes membership filtering efficient when the field has many values or when the query workload is filter-heavy.
schema product {<br> document product {<br> field product_id type int {<br> indexing: attribute | summary<br> attribute: fast-search<br> }<br><br> field title type string {<br> indexing: index | summary<br> }<br><br> field category_ids type array<int> {<br> indexing: attribute | summary<br> attribute: fast-search<br> }<br><br> field category_ancestor_ids type array<int> {<br> indexing: attribute | summary<br> attribute: fast-search<br> }<br><br> field category_slugs type array<string> {<br> indexing: index | summary<br> }<br><br> field category_names type array<string> {<br> indexing: index | summary<br> }<br> }<br>}
Use category_ids when the selected category should match only direct membership. Use category_ancestor_ids when a category page should include all descendants. Many stores use the ancestor field for navigation and retain the direct field for administration, analytics, or exact category reports.
Filter products by category
A category filter can be expressed in a Vespa query using YQL:
select * from product where category_ancestor_ids contains 18
From an application, the value should be inserted as a validated integer rather than concatenated from untrusted input. A typical HTTP request might look like this:
curl --get 'https://search.example.com/search/' \<br> --data-urlencode 'yql=select * from product where category_ancestor_ids contains 18' \<br> --data-urlencode 'query=waterproof boot' \<br> --data-urlencode 'hits=24'
For a direct-membership filter, use the direct field instead:
select * from product where category_ids contains 42
Category filtering should normally be combined with the textual query rather than replacing it. The category condition narrows the candidate set, while the rank profile determines how well each product matches the search terms, availability, price, popularity, or business rules.
Keep category hierarchy synchronized
The WooCommerce REST API exposes product categories through the product category endpoints. A product response commonly includes category objects with an ID, name, and slug. The category endpoint provides the parent relationship needed to construct the complete hierarchy.
A synchronization process should build a category map before indexing products:
- Fetch all active WooCommerce product categories, including their IDs and parent IDs.
- Construct the parent chain for every category.
- Normalize each product’s category list into direct IDs, slugs, names, and ancestor IDs.
- Feed the product document to Vespa using a deterministic document ID such as
product::581.
When a category’s parent changes, every affected product may need to be re-fed because its calculated ancestor list has changed. This is easy to miss if the integration only processes product updates. A reliable agency implementation either reindexes products beneath the changed category or calculates hierarchy expansion during a controlled catalog rebuild.
Use separate category documents when category pages need search
If category pages have their own descriptions, images, merchandising rules, or SEO fields, model them separately:
schema product_category {<br> document product_category {<br> field category_id type int {<br> indexing: attribute | summary<br> attribute: fast-search<br> }<br><br> field name type string {<br> indexing: index | summary<br> }<br><br> field slug type string {<br> indexing: attribute | summary<br> }<br><br> field parent_id type int {<br> indexing: attribute | summary<br> }<br><br> field ancestor_ids type array<int> {<br> indexing: attribute | summary<br> attribute: fast-search<br> }<br><br> field description type string {<br> indexing: index | summary<br> }<br> }<br>}
Feed a category document with a stable ID such as product_category::42. Keep the WooCommerce ID in a field as well, because it allows operational tools to locate documents without depending on the Vespa document ID format.
Category documents can then be searched independently:
select * from product_category where userQuery()
For a category landing page, the application can first retrieve category 42, then issue a product query using category_ancestor_ids contains 42. This keeps category content and product retrieval independent while preserving a straightforward request flow.
Handle deletes, empty categories, and visibility
WooCommerce category synchronization should account for more than newly created terms. When a category is deleted or made inactive, remove or update its category document and update affected products. Do not leave stale category IDs in product documents, because those IDs can continue to produce results even after the category disappears from WooCommerce.
Category membership is also separate from product visibility. A product may belong to a category but still be excluded because it is out of stock, private, unpublished, or outside the current sales channel. Store visibility and inventory fields separately and apply them as query filters or rank constraints.
select * from product where category_ancestor_ids contains 42 and published = true and stock_quantity > 0
Validate the feed with representative catalog cases
Before releasing the integration, test products that have no categories, multiple direct categories, deeply nested categories, renamed slugs, moved parents, and deleted categories. Compare WooCommerce counts with Vespa query results for both direct and descendant-inclusive filters.
Also verify that category IDs are treated as numbers consistently. Mixing numeric IDs with string representations can lead to failed filters or duplicate synchronization logic. Keep the transformation from WooCommerce responses to Vespa fields in one tested component so every indexing path produces the same category shape.