Skip to content

3.2. Autoscaling (APM HPA)

Overview​

Pods are scaled out and in automatically on an APM application metric (TPS). It lets you scale to the real traffic, which CPU and memory alone find hard to catch.

1. How It Works​

APM HPA implements the Kubernetes External Metrics API (external.metrics.k8s.io/v1beta1). When the HPA asks for an external metric, APM HPA queries the OPENMARU APM server for that group's metric and returns the value.

How APM HPA works
  • groupName: the name of the application group in OPENMARU APM.
  • apmAlias: the alias that picks which APM server to ask (the name registered in envHpa at installation). An alias that was never registered is rejected with an error, not silently redirected to another server.

Available metrics​

MetricWhat it isRangetarget.type
tpsRequests handled per second—AverageValue
activeUserActive users—AverageValue
loginUserLogged-in users—AverageValue
avgRTAverage response time (ms)—Value
errorRateError rate0-100Value
apdexDeficit100 - apdex. Higher means worse response0-100Value

Use the target.type from the table. AverageValue divides the value by the pod count. That suits tps, which accumulates per pod, but dividing avgRT, errorRate or apdexDeficit -- which are already averages or ratios -- changes their meaning: a 200 ms response time reads as 50 ms across four pods.

apdex cannot drive autoscaling. It is higher-is-better, while a HorizontalPodAutoscaler can only express "scale up when the target is exceeded" -- so using it directly adds pods when responses are good and removes them when responses degrade. Use apdexDeficit to scale on response quality. apdex itself is read-only.

Fractions arrive in milli-units: a tps of 27.667 comes through as 27667m.

2. Using It — Creating an HPA​

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: tomcat-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: tomcat
minReplicas: 1
maxReplicas: 3
metrics:
- type: External
external:
metric:
name: tps # a metric name from the table above
selector:
matchLabels:
apmAlias: APM-SERVER # which APM server (matches the name in envHpa)
groupName: TOMCAT # the APM application group name
target:
type: AverageValue
value: 50 # scale towards an average of 50 TPS per pod

Applying it:

kubectl apply -f tomcat-hpa.yaml
kubectl get hpa tomcat-hpa

NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
tomcat-hpa Deployment/tomcat 0/50 1 3 1 8s
  • The integration is working once the left-hand value of TARGETS (the current TPS) shows properly.

3. Confirming It Works (a Load Test)​

# generate load
ab -n 18000 -c 10 http://tomcat-openmaru-test1.apps.example.local/

# pods grow once TPS passes the target
kubectl get hpa tomcat-hpa
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
tomcat-hpa Deployment/tomcat 506/50 1 3 3 24m

The external metrics API can also be checked directly.

kubectl get --raw "/apis/external.metrics.k8s.io/v1beta1/namespaces/<namespace>/tps?labelSelector=groupName%3D<group>%2CapmAlias%3D<alias>" | jq .

When the value cannot be fetched, the error message comes through as it is, saying what went wrong.

Error from server: 메트릭 "tps" 를 가져오지 못했다 (그룹 "wrong-name"): ...
Error from server: labelSelector 에 groupName 이 필요하다 (받은 것: "")
Error from server: apmAlias "APM-TYPO" 에 해당하는 환경변수가 없다

3.1 The application group name changes on every redeploy​

groupName is the application group name that APM recognised, and it comes from the OMAPM_APPLICATION_NAME environment variable on the instrumented pod. The default injected by auto-instrumentation contains the ReplicaSet hash.

Pod name helloworld-55cb78b54c-7276g
Group name helloworld-55cb78b54c <- the ReplicaSet hash is in it

Scaling up and down happens inside the same ReplicaSet, so that is fine. But redeploying the Deployment changes the hash, so the group name changes and the HPA's groupName no longer matches.

The HPA then cannot fetch the metric -- and it holds the pod count rather than reducing it. Nothing surfaces the error, so it can go unnoticed for a long time.

To pin it, set OMAPM_APPLICATION_NAME yourself in the pod template. Auto-instrumentation does not overwrite a value that is already there.

spec:
template:
metadata:
labels:
openmaru.io/was-agent: 'true'
spec:
containers:
- name: tomcat
env:
- name: OMAPM_APPLICATION_NAME
value: "tomcat-prod" # unchanged across redeploys

Then use the same name in the HPA.

matchLabels:
groupName: tomcat-prod

To see the group name in use, read the environment variable on an instrumented pod, or try the kubectl get --raw command above and see whether a value comes back.

kubectl get pod <pod> -o jsonpath='{.spec.containers[0].env[?(@.name=="OMAPM_APPLICATION_NAME")].value}'

4. Using Several APM Servers​

With several APM servers, register each under a different name (= apmAlias) in envHpa at installation.

envHpa:
- name: APM-INTERNAL
value: "http://openmaru-apm-server.openmaru-apm.svc.cluster.local:8080"
- name: APM-SERVER
value: "http://192.168.80.190"
- name: APM-INTERNAL_ACCESS_KEY # pair an access key with every alias
value: "<APM API access key>"
- name: APM-SERVER_ACCESS_KEY # that APM server's API access key (<apmAlias>_ACCESS_KEY)
value: "<APM API access key>"

Pair <apmAlias>_ACCESS_KEY with every alias. Without the key the APM server answers HTTP 403 and no metric is fetched.

Then pick between them with matchLabels.apmAlias in the HPA.

matchLabels:
apmAlias: APM-INTERNAL # query the internal APM server
groupName: TOMCAT

If apmAlias matches no name in envHpa, the request is rejected with an error (apmAlias "..." 에 해당하는 환경변수가 없다). It is never redirected to another server, so a typo surfaces immediately -- match the name exactly to use the server you intended.

5. The Key Settings at a Glance​

HPA fieldValueMeaning
metrics[].typeExternalScale on an external metric
metric.nametpsThe metric to use. See the table in section 1
selector.matchLabels.apmAliase.g. APM-SERVERThe alias of the APM server to query (the envHpa name)
selector.matchLabels.groupNamee.g. TOMCATThe APM application group
target.type / target.valueAverageValue / 50The target TPS per pod

6. When It Does Not Work​

SymptomWhat to check
TARGETS shows <unknown>Run kubectl describe hpa <name> and read Events. The reason comes through as a message
Scaling stopped after a redeployThe group name has most likely changed. See 3.1
HTTP 403Check that <apmAlias>_ACCESS_KEY is paired with the alias
apmAlias ... has no matching environment variableWhether apmAlias matches the name in envHpa exactly (case included)
labelSelector needs groupNameselector.matchLabels.groupName is missing
Always reads zeroWhether the group really has traffic and APM is aggregating the metric
Scaling on apdex moves the wrong wayapdex is higher-is-better and cannot be used. Use apdexDeficit
Response time or error rate reads too smalltarget.type was left as AverageValue. Use Value as in the section 1 table
It does not scaleReview minReplicas/maxReplicas and whether target.value is too high