排查 Kubernetes HPA 通过 Prometheus 获取不到 http

部署好了 kube-prometheus 与 k8s-prometheus-adapter （详见之前的博文 k8s 安装 prometheus 过程记录），使用下面的配置文件部署 HPA(Horizontal Pod Autoscaling) 却失败。

apiVersion: autoscaling/v2beta2

kind: HorizontalPodAutoscaler

metadata:

  name: blog-web

spec:

  scaleTargetRef:

    apiVersion: apps/v1

    kind: Deployment

    name: blog-web

  minReplicas: 2

  maxReplicas: 12

  metrics:

    - type: Pods

      pods:

        metric:

          name: http_requests

        target:

          type: AverageValue

          averageValue: 100

错误信息如下：

unable to get metric http_requests: unable to fetch metrics from custom metrics API: the server could not find the metric http_requests for pods

通过下面的命令查看 custom.metrics.k8s.io api 支持的 http_requests（每秒请求数QPS)监控指标：

$kubectl get --raw /apis/custom.metrics.k8s.io/v1beta1/ | jq . | egrep pods/.*http_requests

      "name": "pods/alertmanager_http_requests_in_flight",

      "name": "pods/prometheus_http_requests"

发现只有 prometheus_http_requests 指标，没有所需的 http_requests 开头的指标。

打开 prometheus 控制台，发现 /service-discovery 中没有出现我们想监控的应用 blog-web ，网上查找资料后知道了需要部署 ServiceMonitor 让 prometheus 发现所监控的 service 。

添加下面的 ServiceMonitor 配置文件：

kind: ServiceMonitor

apiVersion: monitoring.coreos.com/v1

metadata:

  name: blog-web-monitor

  labels:

    app: blog-web-monitor

spec:

  selector:

    matchLabels:

      app: blog-web

  endpoints:

  - port: http

部署后还是没有被 prometheus 发现，查看 prometheus 的日志发现下面的错误：

Failed to list *v1.Pod: pods is forbidden: User \"system:serviceaccount:monitoring:prometheus-k8s\" cannot list resource \"pods\" in API group \"\" at the cluster scope

在园子里的博文 PrometheusOperator服务自动发现-监控redis样例中找到了解决方法，将 prometheus-clusterRole.yaml 改为下面的配置：

apiVersion: rbac.authorization.k8s.io/v1

kind: ClusterRole

metadata:

  name: prometheus-k8s

rules:

- apiGroups:

  - ""

  resources:

  - nodes

  - services

  - endpoints

  - pods

  - nodes/proxy

  verbs:

  - get

  - list

  - watch

- apiGroups:

  - ""

  resources:

  - configmaps

  - nodes/metrics

  verbs:

  - get

- nonResourceURLs:

  - /metrics

  verbs:

  - get

重新部署即可

kubectl apply -f prometheus-clusterRole.yaml

注1：如果采用上面的方法还是没被发现，需要强制刷新 prometheus 的配置，参考部署 ServiceMonitor 之后如何让 Prometheus 立即发现。

注2：也可以将 prometheus 配置为自动发现 service 与 pod ，参考园子里的博文 prometheus配置pod和svc的自动发现和监控与 PrometheusOperator服务自动发现-监控redis样例。

但是这时还有问题，虽然 service 被 prometheus 发现了，但 service 所对应的 pod 一个都没被发现。

production/blog-web-monitor/0 (0/19 active targets)

排查后发现是因为 ServiceMonitor 与 Service 配置不对应，Service 配置文件中缺少 ServiceMonitor 配置中 matchLabels 所对应的 label ，ServiceMonitor 中的 port 没有对应 Service 中的 ports 配置，修正后的配置如下：

service-blog-web.yaml

apiVersion: v1

kind: Service

metadata:

  name: blog-web

  labels:

    app: blog-web

spec:

  type: NodePort

  selector:

    app: blog-web

  ports:

  - name: http-blog-web

    nodePort: 30080

    port: 80

    targetPort: 80

servicemonitor-blog-web.yaml

kind: ServiceMonitor

apiVersion: monitoring.coreos.com/v1

metadata:

  name: blog-web-monitor

  labels:

    app: blog-web

spec:

  selector:

    matchLabels:

      app: blog-web

  endpoints:

  - port: http-blog-web

用修正后的配置部署后，pod 终于被发现了：

production/blog-web-monitor/0 (0/5 up)

但是这些 pod 全部处于 down 状态。

Endpoint	                      State	 Scrape Duration	Error

http://192.168.107.233:80/metrics DOWN	 server returned HTTP status 400 Bad Request

通过园子里的博文使用Kubernetes演示金丝雀发布知道了原来需要应用自己提供 metrics 监控指标数据让 prometheus 抓取。

标准Tomcat自带的应用没有/metrics这个路径，prometheus获取不到它能识别的格式数据，而指标数据就是从/metrics这里获取的。所以我们使用标准Tomcat不行或者你就算有这个/metrics这个路径，但是返回的格式不符合prometheus的规范也是不行的。

我们的应用是用 ASP.NET Core 开发的，所以选用了 prometheus-net ，由它提供 metrics 数据给 prometheus 抓取。

安装 nuget 包

dotnet add package prometheus-net.AspNetCore

添加 HttpMetrics 中间件

app.UseRouting();

app.UseHttpMetrics();

添加 MapMetric 路由

app.UseEndpoints(endpoints =>

{

   endpoints.MapMetrics();

};

当通过下面的命令确认通过 /metrics 路径可以获取监控数据时，

$ docker exec -t $(docker ps -f name=blog-web_blog-web -q | head -1) curl 127.0.0.1/metrics | grep http_request_duration_seconds_sum

http_request_duration_seconds_sum{code="200",method="GET",controller="AggSite",action="SiteHome"} 0.44973779999999997

http_request_duration_seconds_sum{code="200",method="GET",controller="",action=""} 0.0631272

Prometheus 控制台 /targets 页面就能看到 blog-web 对应的 pod 都处于 up 状态。

production/blog-web-monitor/0 (5/5 up)

这时通过 custom metrics api 可以查询到一些 http_requests 相关的指标。

$ kubectl get --raw /apis/custom.metrics.k8s.io/v1beta1/ | jq . | egrep pods/*/http_requests

      "name": "pods/http_requests_in_progress",

      "name": "pods/http_requests_received"

这里的 http_requests_received 就是 QPS（每秒请求数）指标数据，用下面的命令请求 custom metrics api 获取数据：

kubectl get --raw /apis/custom.metrics.k8s.io/v1beta1/namespaces/production/pods/*/http_requests_received | jq .

其中1个 pod 的 http_requests_received 指标数据如下：

{

  "kind": "MetricValueList",

  "apiVersion": "custom.metrics.k8s.io/v1beta1",

  "metadata": {

    "selfLink": "/apis/custom.metrics.k8s.io/v1beta1/namespaces/production/pods/%2A/http_requests_received"

  },

  "items": [

    {

      "describedObject": {

        "kind": "Pod",

        "namespace": "production",

        "name": "blog-web-65f7bdc996-8qp5c",

        "apiVersion": "/v1"

      },

      "metricName": "http_requests_received",

      "timestamp": "2020-01-18T14:35:34Z",

      "value": "133m",

      "selector": null

    }

  ]

}

其中的 133m 表示 0.133 。

然后就可以在 HPA 配置文件中基于这个指标进行自动伸缩

apiVersion: autoscaling/v2beta2

kind: HorizontalPodAutoscaler

metadata:

  name: blog-web

spec:

  scaleTargetRef:

    apiVersion: apps/v1

    kind: Deployment

    name: blog-web

  minReplicas: 5

  maxReplicas: 12

  metrics:

  - type: Pods

    pods:

      metric:

        name: http_requests_received

      target:

        type: AverageValue

        averageValue: 100

终于搞定了！

# kubectl get hpa

NAME       REFERENCE             TARGETS    MINPODS   MAXPODS   REPLICAS   AGE

blog-web   Deployment/blog-web   133m/100   5         12        5          4d

排查 Kubernetes HPA 通过 Prometheus 获取不到 http_requests 指标的问题的更多相关文章

终于成功部署 Kubernetes HPA 基于 QPS 进行自动伸缩
昨天晚上通过压测验证了 HPA 部署成功了. 所使用的 HPA 配置文件如下: apiVersion: autoscaling/v2beta2 kind: HorizontalPodAutoscale ...
kubernetes之监控Prometheus实战--prometheus介绍--获取监控（一）
Prometheus介绍 Prometheus是一个最初在SoundCloud上构建的开源监控系统 .它现在是一个独立的开源项目,为了强调这一点,并说明项目的治理结构,Prometheus 于2016 ...
Kubernetes 监控：Prometheus Adpater =》自定义指标扩缩容
使用 Kubernetes 进行容器编排的主要优点之一是,它可以非常轻松地对我们的应用程序进行水平扩展.Pod 水平自动缩放(HPA)可以根据 CPU 和内存使用量来扩展应用,前面讲解的 HPA 章节 ...
Kubernetes 监控：Prometheus Operator + Thanos ---实践篇
具体参考网址:https://www.cnblogs.com/sanduzxcvbnm/p/16291296.html 本章用到的yaml文件地址:https://files.cnblogs.com/ ...
Kubernetes HPA 使用详解
文章转载自:https://www.qikqiak.com/post/k8s-hpa-usage/ Kubernetes 提供了这样的一个资源对象:Horizontal Pod Autoscaling ...
Kubernetes 监控：Prometheus Operator
安装前面的章节中我们学习了用自定义的方式来对 Kubernetes 集群进行监控,基本上也能够完成监控报警的需求了.但实际上对上 Kubernetes 来说,还有更简单方式来监控报警,那就是 Pro ...
kubernetes学习笔记之十二：资源指标API及自定义指标API
第一章.前言以前是用heapster来收集资源指标才能看,现在heapster要废弃了从1.8以后引入了资源api指标监视资源指标:metrics-server(核心指标) 自定义指标:prome ...
Kubernetes之利用prometheus监控K8S集群
prometheus它是一个主动拉取的数据库,在K8S中应该展示图形的grafana数据实例化要保存下来,使用分布式文件系统加动态PV,但是在本测试环境中使用本地磁盘,安装采集数据的agent使用Da ...
在Kubernetes下部署Prometheus
使用ConfigMaps管理应用配置当使用Deployment管理和部署应用程序时,用户可以方便了对应用进行扩容或者缩容,从而产生多个Pod实例.为了能够统一管理这些Pod的配置信息,在Kuber ...

随机推荐

php部署后错误排查流程
未使用框架的php程序不可用时,没有框架提供的调试信息,因此要按照请求的整个生命周期来调试程序, 具体错误依次排查网络,服务器,环境,代码的步骤层层深入,最终定位到错误的发生点. 1 访问程序部署的服 ...
在 Vue 中使用 Typescript
前言恕我直言,用 Typescript 写 Vue 真的很难受,Vue 对 ts 的支持一般,如非万不得已还是别在 Vue 里边用吧,不过听说 Vue3 会增强对 ts 的支持,正式登场之前还是期待 ...
2次方的期望dp
某一天WJMZBMR在打osu~~~但是他太弱逼了,有些地方完全靠运气:( 我们来简化一下这个游戏的规则有n次点击要做,成功了就是o,失败了就是x,分数是按comb计算的,连续a个com ...
python中常⽤的excel模块库
python中常用的excel模块库&安装方法 openpyxl openpyxl是⼀个Python库,用于读取/写⼊Excel 2010 xlsx / xlsm / xltx / xltm⽂ ...
AVR单片机教程——串口接收
本文隶属于AVR单片机教程系列. 上一讲中,我们实现了单片机开发板向电脑传输数据.在这一讲中,我们将通过电脑向单片机发送指令,让单片机根据指令控制LED.这一次,两端的TX与RX需要交叉连接,单片 ...
Linux.vim.多行复制、删除、剪切
复制: //单行复制+粘贴 yy + p:复制光标所处当前行, 敲p粘贴在光标处. //多行复制+粘贴 n + yy + p:复制光标所在行起以下n行(含当前行), 敲yy复制光标所处当前行, 敲p粘 ...
restframework 认证、权限、频率组件
一.认证 1.表的关系 class User(models.Model): name = models.CharField(max_length=32) pwd = models.CharField( ...
Scala 学习（4）之「类——基本概念2」
目录内部类 extends override和super override field isInstanceOf和asInstanceOf getClass和classOf 内部类 import s ...
VSCODE更改文件时，提示EACCES permission denied的解决办法(mac电脑系统)
permission denied:权限问题具体解决办法: 1.在项目文件夹右键-显示简介-点击右下角解锁 2.权限全部设置为读与写 3.最关键一步:点击"应用到包含的项目",这 ...
Windows环境下配置robotframework
Robot Framework安装准备一.python3.6以上版本安装过程中勾选“add python to path”,就可以自动配置好环境变量. 安装完成后在命令行输入python,如下图所 ...

排查 Kubernetes HPA 通过 Prometheus 获取不到 http_requests 指标的问题

排查 Kubernetes HPA 通过 Prometheus 获取不到 http_requests 指标的问题的更多相关文章

随机推荐

热门专题