elasticsearch - fulltext search for words with special/reserved characters

Question

I am indexing documents that may contain any special/reserved characters in their fulltext body. For example "PDF/A is an ISO-standardized version of the Portable Document Format..."

I would like to be able to search for pdf/a without having to escape the forward slash.

How should i analyze my query-string and what type of query should i use?

Can you share what you have tried? What do your mappings and query look like as a starting point? — eemp

eemp eemp · Accepted Answer · 2016-09-06T19:03:37

The default standard analyzer will tokenize a string like that so that "PDF" and "A" are separate tokens. The "A" token might get cut out by the stop token filter (See Standard Analyzer). So without any custom analyzers, you will typically get any documents with just "PDF".

You can try creating your own analyzer modeled off the standard analyzer that includes a Mapping Char Filter. The idea would that "PDF/A" might get transformed into something like "pdf_a" at index and query time. A simple match query will work just fine. But this is a very simplistic approach and you might want to consider how '/' characters are used in your content and use slightly more complex regex filters which are also not perfect solutions.

Sorry, I completely missed your point about having to escape the character. Can you elaborate on your use case if this turns out to not be helpful at all?

elasticsearch - fulltext search for words with special/reserved characters

2 Answers