Define the product document before indexing
A reliable WooCommerce search experience starts with a deliberate Elasticsearch document structure. Do not send the entire wp_posts row or raw post metadata and expect Elasticsearch to infer the right behavior. Build one search document per product, or per variation when customers need to search and filter variations independently.
A typical product document can contain separate fields for the values customers search as text and the values they match exactly:
{
"product_id": 1842,
"title": "Organic Cotton Crew Neck T-Shirt",
"sku": "TSH-COT-ORG-BLK-M",
"description": "A lightweight crew neck T-shirt made from certified organic cotton.",
"short_description": "Certified organic cotton T-shirt.",
"status": "publish",
"stock_status": "instock",
"categories": ["T-Shirts", "Organic Clothing"]
}
Keep the source fields in the document even if you also create a combined search field. Source fields are useful for result display, debugging, reindexing, and future query changes.
Map titles and descriptions as searchable text
Product titles and descriptions should normally use the text type. Elasticsearch analyzes these fields into terms, allowing a query for organic cotton to match a title such as Organic Cotton Crew Neck T-Shirt.
SKUs have different requirements. They are identifiers, not natural language. Map the primary SKU as keyword so the complete value can be matched, filtered, sorted, and aggregated without tokenization.
A practical Elasticsearch mapping is:
{
"mappings": {
"properties": {
"product_id": { "type": "long" },
"title": {
"type": "text",
"fields": {
"keyword": { "type": "keyword", "ignore_above": 256 }
}
},
"sku": { "type": "keyword", "normalizer": "sku_normalizer" },
"description": { "type": "text" },
"short_description": { "type": "text" },
"status": { "type": "keyword" },
"stock_status": { "type": "keyword" },
"categories": { "type": "keyword" }
}
},
"settings": {
"analysis": {
"normalizer": {
"sku_normalizer": {
"type": "custom",
"filter": ["lowercase", "asciifolding"]
}
}
}
}
}
The title.keyword multi-field is useful when an integration needs an unanalyzed title for sorting or exact filtering. It does not replace the analyzed title field used for normal full-text search.
The SKU normalizer makes TSH-COT-ORG-BLK-M and tsh-cot-org-blk-m equivalent for exact matching while preserving the complete identifier as one term. A normalizer can apply character-level operations such as lowercasing, but it cannot tokenize the value like a text analyzer.
Clean WooCommerce content before sending it
WooCommerce descriptions commonly contain HTML, shortcodes, embedded blocks, and presentation markup. Indexing that content unchanged can produce noisy search results and expose implementation details in snippets.
Before indexing a description:
- Load the current product content and metadata.
- Remove shortcodes and irrelevant markup.
- Convert meaningful HTML structure to readable text.
- Decode entities and normalize whitespace.
- Remove duplicated content where the short description is repeated in the full description.
- Exclude private, draft, trashed, and catalog-hidden products according to the store’s visibility rules.
For example, this WooCommerce content:
<p>Lightweight <strong>organic cotton</strong> fabric.</p>
[product_table id="1842"]
should become a searchable value similar to:
Lightweight organic cotton fabric.
Use WordPress sanitization and content filters consistently with the storefront, but avoid blindly indexing the output of the_content if that filter adds widgets, related products, or other non-product text. A dedicated indexing normalizer is usually more predictable.
Handle SKUs as identifiers
A product SKU should be indexed as a single exact value. Avoid mapping it only as text, because analysis may split values containing hyphens, slashes, or spaces. A search for a partial token could then return products that do not have the requested SKU.
For an exact SKU lookup, use a term query against the keyword field:
{
"query": {
"term": {
"sku": "tsh-cot-org-blk-m"
}
}
}
Because the normalizer lowercases the value at index and query time, users can enter the SKU in uppercase or lowercase. If the business requires prefix searches such as TSH-COT, add a dedicated search strategy rather than changing the primary SKU field to analyzed text. Options include a prefix query, an edge_ngram field, or a separate autocomplete field, depending on catalog size and query volume.
If variations have their own SKUs, decide whether a parent product document should contain an array of variation SKUs or whether each variation should be indexed separately. For example:
{
"product_id": 1842,
"variation_skus": ["TSH-COT-ORG-BLK-M", "TSH-COT-ORG-BLK-L"]
}
An array of keyword values supports exact matching, but separate variation documents are usually better when the result must show the matching size, color, price, or stock status.
Search across title, SKU, and descriptions
Use a multi_match query when a customer enters normal product text. Give the title more weight than descriptions, and search the SKU field separately so identifier matches receive an appropriate score.
{
"query": {
"bool": {
"should": [
{
"multi_match": {
"query": "organic cotton",
"fields": ["title^4", "short_description^2", "description"]
}
},
{
"term": {
"sku": "organic cotton"
}
}
],
"minimum_should_match": 1,
"filter": [
{ "term": { "status": "publish" } },
{ "term": { "stock_status": "instock" } }
]
}
}
}
In application code, do not send the user’s input to the SKU term clause without normalizing it in the same way as the index. For a general search term such as organic cotton, the SKU clause may not produce a match, while the title and description clauses still work. For a known SKU lookup, use a dedicated exact-match path and return the product directly when possible.
A match_phrase clause can improve title relevance for multi-word queries:
{
"match_phrase": {
"title": {
"query": "organic cotton",
"boost": 6
}
}
}
Use this as an additional scoring clause rather than replacing the normal match or multi_match query. Phrase matching alone can be too restrictive for shoppers who type terms in a different order.
Build and update documents from WooCommerce events
Indexing should be driven by product lifecycle events instead of only by scheduled full imports. At minimum, handle product creation, product updates, status changes, deletions, stock changes, price changes, and variation updates.
A WordPress integration can register actions such as save_post_product and variation-related product hooks, but it should avoid indexing during autosaves, revisions, imports that are not complete, or recursive saves. The indexing handler should enqueue a product ID rather than performing a slow Elasticsearch request inside the page request.
A queue worker can then:
- Load the latest product state from WooCommerce.
- Rebuild the complete Elasticsearch document.
- Send an index request using the stable product ID.
- Remove the document when the product is deleted or no longer eligible for search.
- Record failures for retry and operational review.
Rebuilding the complete document is safer than issuing partial updates for every metadata change. It prevents stale titles, descriptions, SKUs, and visibility flags from remaining in the index when one field changes.
For a product update, use an idempotent request such as:
PUT /products/_doc/1842
Content-Type: application/json
{
"product_id": 1842,
"title": "Organic Cotton Crew Neck T-Shirt",
"sku": "tsh-cot-org-blk-m",
"description": "A lightweight crew neck T-shirt made from certified organic cotton.",
"status": "publish",
"stock_status": "instock"
}
The document ID should remain stable. Do not use a random ID for each update, or old versions will remain searchable.
Use bulk indexing for initial imports
For a new catalog or a complete reindex, use the Elasticsearch Bulk API rather than making one HTTP request per product. Generate alternating action and document lines:
{ "index": { "_index": "products-v1", "_id": "1842" } }
{ "product_id": 1842, "title": "Organic Cotton Crew Neck T-Shirt", "sku": "tsh-cot-org-blk-m", "description": "A lightweight crew neck T-shirt made from certified organic cotton." }
Send batches sized according to document size and cluster capacity, inspect the errors property in every bulk response, and retry only failed items. A successful HTTP response does not mean every document was indexed successfully.
For large WooCommerce catalogs, create a versioned index such as products-v1, bulk index the catalog, run representative search checks, and then update an alias from the old index to the new one. This avoids exposing a partially rebuilt index to shoppers and makes rollback possible.
Verify the indexed values
When a product does not appear in search, inspect the actual indexed document and mapping before changing the query. Confirm that:
- The product is published and eligible for the catalog.
- The title and descriptions contain cleaned text rather than empty HTML or shortcodes.
- The SKU is present and normalized as expected.
- The document ID is stable and not duplicated.
- The field is mapped as
textorkeywordfor its intended use. - The index alias used by the application points to the current index.
- Bulk responses did not contain item-level failures.
Use the Analyze API to check how a title or SKU is processed, and use termvectors or a test query to confirm which terms are actually searchable. This is especially useful after changing analyzers or migrating an existing WooCommerce integration to a new index mapping.