ElasticSearch 7.x现网运行问题汇集3

最新推荐文章于 2024-08-18 22:18:25 发布

旻璿gg

最新推荐文章于 2024-08-18 22:18:25 发布

阅读量947

点赞数 13

分类专栏： Elastic Search 大数据文章标签： elasticsearch 大数据

本文链接：https://blog.csdn.net/sencloud/article/details/135726486

版权

大数据同时被 2 个专栏收录

13 篇文章 0 订阅

订阅专栏

Elastic Search

3 篇文章 0 订阅

订阅专栏

文章讲述了在现网ElasticSearch中遇到的故障，涉及unassigned_shards数量不减的问题。通过对节点状态、健康检查、磁盘使用情况的分析，发现原因是磁盘空间不足。解决方案包括设置临时关闭分配、扩容磁盘、重启Elasticsearch集群，并逐步恢复分配策略。

摘要由CSDN通过智能技术生成

问题描述

某现网ElasticSearch 故障，很长时间unassgined_shards的数量都不减少。

原因分析与解决方案：

先了解整体状态，使用Postman请求，如下几个请求命令：

GET /_cat/indices
GET /_cat/shards
GET /_cluster/health
GET /_cat/nodes?v
GET /_cat/health?v
GET /_cluster/allocation/explain
POST /_cluster/reroute?retry_failed=true

恢复了部分，但是还是有shards没恢复，取回/_cluster/allocation/expain的response，才发现日志显示：

"disk_threshold","the node is above the low watermark cluster setting [cluster.routing.allocation.disk.watermark.low=85%], using more disk space than the maximum allowed [85.0%], actual free: [12.239612269812415%]"

确认了分片无法指向的原因是节点磁盘使用率超过85%，即安排磁盘扩容，然后再重启ES集群解决。具体操作重启步骤：

第一步：PUT /_cluster/settings
Body里的内容：

{
  "transient": {
    "cluster.routing.allocation.enable": "none"
  }
}

第二步：
systemctl stop elasticsearch或kill {es的pid}，注意不是kill -9
这时候要等，通过ps -ef | grep elasticsearch看进程结束没。
进程结束后，再进入第三步。

第三步：
systemctl start elasticsearch或su - esuser进入elasticsearch的bin目录，执行./elasticsearch -d命令

观察es的日志，直到它加入集群，再重启下一台。

重复2、3两步，全部节点重启完成后执行
第四步：

PUT  /_cluster/settings
{
   "transient" : {
       "cluster.routing.allocation.enable" : "all"
   }
 }

旻璿gg

关注

13
点赞
踩
5

收藏

觉得还不错? 一键收藏
打赏
0
评论
复制链接

分享到 QQ

分享到新浪微博

扫一扫

专栏目录