二十一 _validate验证、结果排序规则、String排序问题、score计算方式

最新推荐文章于 2023-07-01 23:11:50 发布

tianlan996

最新推荐文章于 2023-07-01 23:11:50 发布

阅读量543

点赞数

分类专栏： Elasticsearch核心知识

本文链接：https://blog.csdn.net/tianlan996/article/details/93914429

版权

Elasticsearch核心知识专栏收录该内容

31 篇文章 0 订阅

订阅专栏

1. _validate可以用来验证搜索是否合法。

示例：

GET /test_index/test_type/_validate/query?explain
{
"query": {
"math": {
"test_field": "test"
}
}
}

结果：

{
"valid": false,
"error": "org.elasticsearch.common.ParsingException: no [query] registered for [math]"
}

2. 结果排序规则：

默认情况下，是按照_score降序排序的

然而，某些情况下，可能没有用到_score，比如说filter

GET /_search
{
"query" : {
"bool" : {
"filter" : {
"term" : {
"author_id" : 1
}
}
}
}
}

当然，也可以是constant_score

GET /_search
{
"query" : {
"constant_score" : {
"filter" : {
"term" : {
"author_id" : 1
}
}
}
}
}

定制排序规则：

GET /company/employee/_search
{
"query": {
"constant_score": {
"filter": {
"range": {
"age": {
"gte": 30
}
}
}
}
},
"sort": [ --- 制定排序规则
{
"join_date": {
"order": "asc"
}
}
]
}

3. String排序问题

如果对一个string field进行排序，结果往往不准确，因为分词后是多个单词，再排序就不是我们想要的结果了

通常解决方案是，将一个string field建立两次索引，一个分词，用来进行搜索；一个不分词，用来进行排序

PUT /website
{
"mappings": {
"article": {
"properties": {
"title": {
"type": "text",
"fields": {
"raw": {
"type": "string",
"index": "not_analyzed"
}
},
"fielddata": true
},
"content": {
"type": "text"
},
"post_date": {
"type": "date"
},
"author_id": {
"type": "long"
}
}
}
}
}

GET /website/article/_search
{
"query": {
"match_all": {}
},
"sort": [
{
"title.raw": { ---使用未分词的字段排序
"order": "desc"
}
}
]
}

4. score是如何被计算出来的

relevance score算法，简单来说，就是计算出，一个索引中的文本，与搜索文本，他们之间的关联匹配程度

Elasticsearch使用的是 term frequency/inverse document frequency算法，简称为TF/IDF算法

Term frequency：搜索文本中的各个词条在field文本中出现了多少次，出现次数越多，就越相关

搜索请求：hello world

doc1：hello you, and world is very good
doc2：hello, how are you

Inverse document frequency：搜索文本中的各个词条在整个索引的所有文档中出现了多少次，出现的次数越多，就越不相关

搜索请求：hello world

doc1：hello, today is very good
doc2：hi world, how are you

比如说，在index中有1万条document，hello这个单词在所有的document中，一共出现了1000次；world这个单词在所有的document中，一共出现了100次

doc2更相关

Field-length norm：field长度，field越长，相关度越弱

搜索请求：hello world

doc1：{ "title": "hello article", "content": "babaaba 1万个单词" }
doc2：{ "title": "my article", "content": "blablabala 1万个单词，hi world" }

hello world在整个index中出现的次数是一样多的

doc1更相关，title field更短

分析score是如何计算出来的接口：

GET /test_index/test_type/_search?explain
{
"query": {
"match": {
"test_field": "test hello"
}
}
}

分析一个document是如何被匹配上的：

GET /test_index/test_type/6/_explain
{
"query": {
"match": {
"test_field": "test hello"
}
}
}

tianlan996

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
二十一 _validate验证、结果排序规则、String排序问题、score计算方式

1. _validate可以用来验证搜索是否合法。示例：GET /test_index/test_type/_validate/query?explain{ "query": { "math": { "test_field": "test" } }}结果：{ "valid": false, "error": "org.elastics...
复制链接

扫一扫