0
votes

My use case is like this: for a query iphone charger, I am getting higher relevance for results, having name, iphone charger coupons than with name iphone charger, possibly because of better match in description and other fields. Boosting name field isn't helping much unless I skew the importance drastically. what I really need is tf/idf boost within name field

to quote elasticsearch blog:

the frequency of a term in a field is offset by the length of the field. However, the practical scoring function treats all fields in the same way. It will treat all title fields (because they are short) as more important than all body fields (because they are long).

I need to boost this more important value for a particular field. Can we do this with function score or any other way?

1
I think you are thinking about this from the wrong angle. Think about what are your requirements. "iphone charger" is shorter than the other text. Is this the rule by which you want to favor this text? You need to think about the matching rules and then think about how you can achieve this. Messing up with tf/idf is not desired, imo. - Andrei Stefan
@AndreiStefan for me the name field is naturally more important than any other field. But within name, even after match relevance is not quite acceptable because one extra word changes everything. So less words + more match gives me desired result - device_exec
Have you seen this? - Andrei Stefan
yes, ignoring tf/idf may lead me reverse way, causing products with even larger name fields to show up for simple iphone charger - device_exec

1 Answers

0
votes

A one term difference in length is not much of a difference to the scoring algorithm (and, in fact, can vanish entirely due to imprecision on the length norm). If there are hits on other fields, you have a lot of scoring elements to fight against.

A dis_max would probably be a reasonable approach to this. Instead of all the additive scores and coords and such you are trying to overcome, it will simply select the score of the best matching subquery. If you boost the query against title, you can ensure matches there are strongly preferred.

You can then assign a "tie_breaker", so that the score against the description subquery is factored in only when "title" scores are tied.

{
    "dis_max" : {
        "tie_breaker" : 0.2,
        "queries" : [
            {
                "terms" : { 
                    "age" : ["iphone", "charger"],
                    "boost" : 10
                }
            },
            {
                "terms" : {
                    "description" : ["iphone", "charger"]
                }
            }
        ]
    }
}

Another approach to this sort of thing, if you absolutely know when you have an exact match against the entire field, is to separately index an untokenized version of that field, and query that field as well. Any match against the untokenized version of the field will be an exact match again the entire field contents. This would prevent you needing to relying on the length norm to make that determination.