R7. PromQL Integration Reference
Diátaxis: Reference · Audience: operators / administrators (with PromQL or Grafana experience) ← Back to contents
If you are already comfortable with Grafana dashboards or PromQL, the WAS, JVM, system, WEB, and DBMS metrics OPENMARU APM collects can be queried directly with standard PromQL-compatible queries. Collection, storage, querying, and visualization all happen in APM alone, with no separate Prometheus server, exporter, or scrape configuration, and existing Grafana dashboard assets can be reused.
- APM's fine-grained metrics as they are -- reach JVM, GC, thread, DB connection pool, and transaction metrics with standard PromQL
- Rich labels -- filter and aggregate precisely with APM context such as
instance_id,ip_addr, andagent_type - rate/increase follow how a metric is stored -- a metric stored as a delta is summed over the window, and a metric already stored as a per-second value is averaged over the window
- Relative time expressions -- intuitive forms such as
nowandnow-1hare supported (a convenience standard Prometheus does not have)
For what each metric means, see R2. Chart Metric Reference and E3. What the Metrics Mean. To gather charts inside the console, H3. My Dashboard is enough -- PromQL is the tool for external integration and automation.
How to open it -- in a browser, http://{apm-server}/monitoring/api/v1/metric-explorer
(the Metric Explorer UI). Sign in with an API access key -- go to the left menu ▸ Settings ▸
Users, press the edit user button, and create and copy the key.
Getting Started in 5 Minutes
The simplest query -- one line gives the JVM heap usage (current value) of every WAS instance.
jvm_heap_heapUsed
To see only a particular instance, or only a particular IP range (regular expression):
jvm_heap_heapUsed{instance_id="apm-was-01"}
jvm_heap_heapUsed{ip_addr=~"192\\.168\\.80\\..*"}
Apply time functions to get the maximum over the last hour and the 5-minute average GC frequency:
max_over_time(jvm_heap_heapUsed[1h])
rate(gc_G1_Young_Generation_gcCount[5m])
Note Choose the exact metric name from the Metric Explorer dropdown, or check it with
GET /api/v1/label/__name__/values(about 250 of them). GC metrics are namedgc_<collector>_gcCountandgc_<collector>_gcTime; the collector name depends on the JVM GC setting (for examplegc_G1_Young_Generation_gcCount,gc_PS_Scavenge_gcCount). The examples in this document use G1.
Querying a Specified Period
A range query (/api/v1/query_range) takes the start and end times and a step (the data interval).
Four time formats are supported.
| Format | Example | Note |
|---|---|---|
| Relative time | now, now-1h, now-30m, now-1d | The most convenient -- units s/m/h/d |
| Unix epoch (seconds) | 1704067200 | shell date +%s |
| Unix epoch with a fraction | 1704067200.123 | Millisecond precision |
| RFC3339 | 2026-05-10T09:00:00+09:00 | Can include a time zone |
Frequently used call examples:
# The last hour -- the most common case (step automatically 15s)
curl -G "http://{apm-server}/monitoring/api/v1/query_range" \
--data-urlencode "query=jvm_heap_heapUsed" \
--data-urlencode "start=now-1h" \
--data-urlencode "end=now"
# The weekly trend over the last 7 days (step automatically 5m)
curl -G "http://{apm-server}/monitoring/api/v1/query_range" \
--data-urlencode "query=max_over_time(cpu_usage_cpuUsage[1h])" \
--data-urlencode "start=now-7d" \
--data-urlencode "end=now"
# Specifying the step directly (15-second interval)
curl -G "http://{apm-server}/monitoring/api/v1/query_range" \
--data-urlencode "query=rate(gc_G1_Young_Generation_gcCount[5m])" \
--data-urlencode "start=now-30m" \
--data-urlencode "end=now" \
--data-urlencode "step=15s"
# A particular date range (KST)
curl -G "http://{apm-server}/monitoring/api/v1/query_range" \
--data-urlencode "query=jvm_heap_heapUsed{agent_type=\"WAS\"}" \
--data-urlencode "start=2026-05-10T09:00:00+09:00" \
--data-urlencode "end=2026-05-11T18:00:00+09:00"
Without a step, it is calculated automatically from the query range:
| Query range | Automatic step | Data points (example) | Suggested use |
|---|---|---|---|
| 12 hours or less | 15s | 1 hour ≈ 240 | Real time · standard dashboards |
| 12 hours to 2 days | 1m | 24 hours ≈ 1,440 | Day-to-day comparison |
| 2 to 8 days | 5m | 7 days ≈ 2,016 | Weekly trend |
| 8 to 21 days | 10m | Fortnightly trend | |
| 21 days or more | 30m | 30 days ≈ 1,440 | Monthly comparison |
Note The raw data is stored every 2 seconds. A smaller step is denser but loads the server more -- start with a large step and reduce it when you need to.
Range queries follow these rules.
- The timestamps in the response are aligned to step boundaries. When
startis 09:00:30 and the step is 1m, the first point is 09:00:00. This is because the storage groups values by step boundaries. - If
endis beforestart, you get abad_dataerror (HTTP 400). - If (
end−start) ÷ step is more than 11,000, you get abad_dataerror (HTTP 400). Specify a larger step. When the step is omitted, it is increased so that it stays within this limit (ranges longer than about 229 days).
The amount one query reads from the storage is also limited. The same limits apply to instant queries.
- One selector or window function can read up to 2,000,000 rows (series × rows per series). The rows per series are the number of steps for a selector without a function, the number of small windows for a window function, and 1 for an instant query.
So the number of series it can read is 2,000,000 ÷ rows per series. For example, a 3-day range with a 30-second step has 8,641 rows per series, so up to 231 series.
Beyond that, the query gets a
bad_dataerror (HTTP 400)query selects too many series (more than <series limit> for this range and step); use a coarser step, a shorter range or a narrower selector. The storage is asked for at most one series more than the limit, so most of the extra series are rejected before they are read. This request count applies per storage unit (measurement), so a metric split across several measurements can read that many more. Use a larger step, shorten the query period, or narrow the series with labels. You can change the limit with the collection server system propertypromql.leaf.max.rows. Calculations that read raw samples (the raw sample window functions, therate,irate, andincreasecalculated from raw samples in "How It Differs from Prometheus" below, andtimestamp()) use the raw sample limit instead of the limits in this section (see "Supported Functions" below). - All selectors and window functions of one query can read up to 8,000,000 rows in total. Beyond that, the query gets a
bad_dataerror (HTTP 400)query reads too many rows (<rows read> > <limit>); use a coarser step, a shorter range or a narrower selector. You can change the limit with the collection server system propertypromql.query.max.rows. - Selectors and window functions with more than 5,000 rows per series (for example, a 3-day range with a 30-second step) read the storage at most 4 at a time across the whole collection server. When 4 are already reading, the next one waits up to 10 seconds.
If its turn still does not come, the query gets a
timeouterror (HTTP 503)too many concurrent heavy queries; retry later. Run the query again a little later. Queries with 5,000 rows per series or fewer (such as the usual dashboard panels) do not wait. You can change the threshold, the number at a time, and the wait time with the collection server system propertiespromql.heavy.min.rows.per.series,promql.heavy.max.concurrent, andpromql.heavy.wait.ms(milliseconds).promql.heavy.max.concurrentis read once when the collection server starts.
If you need only one value at a single point in time, use an instant query (/api/v1/query):
# The heap usage now / one hour ago
curl -G "http://{apm-server}/monitoring/api/v1/query" \
--data-urlencode "query=jvm_heap_heapUsed"
curl -G "http://{apm-server}/monitoring/api/v1/query" \
--data-urlencode "query=jvm_heap_heapUsed" \
--data-urlencode "time=now-1h"
Metric Names and Labels
APM metrics are distinguished by five label dimensions. They are used directly in query filters and aggregation.
| Label | Meaning | Example values |
|---|---|---|
__name__ | The metric name -- in the form {nameSpace}_{metricName}_{field} | jvm_heap_heapUsed, cpu_usage_cpuUsage |
ip_addr | The IP of the host the agent is installed on | 192.168.80.190 |
agent_type | The agent type | WAS, SYS, DBMS, WEB |
instance_id | The instance identifier (unique within the host) | apm-was-01, nginx-1 |
namespace | The metric group | jvm, transaction, cpu, memory |
Prometheus-standard labels are generated automatically as well -- instance
(= {ip_addr}:{instance_id}) and job (= agent_type).
The app_name label is the application name of the instance. A WAS instance gets the
application name its agent reports; the name of a custom group created by a user is never
used. To filter by a custom group, use a matcher such as metric{app_name="<group name>"}.
For metric names, only one form has to be known: {nameSpace}_{metricName}_{field} (the
trailing _field may be omitted -- omitting it selects the default field).
| Example | Meaning |
|---|---|
jvm_heap_heapUsed | The used field of the JVM heap |
jvm_heap_heapMax | The maximum field of the same metric |
cpu_usage_cpuUsage | CPU utilization |
cpu_usage_systemCpuUsage | System-wide CPU utilization |
All four standard label matchers are supported, and several conditions are combined with commas:
metric{label="value"} # equal
metric{label!="value"} # not equal
metric{label=~"regex"} # regular expression match
metric{label!~"regex"} # regular expression non-match
jvm_heap_heapUsed{ip_addr="192.168.80.190", instance_id=~"apm-was-.*", agent_type="WAS"}
Regular expressions use the same RE2 syntax as Prometheus (Go regexp) and must match the full value. POSIX character classes such as [[:alpha:]], named groups (?P<name>…), and the flags (?i), (?m), (?s) are accepted. The following syntax, which RE2 does not accept, returns a 400 error (invalid regular expression …: <reason>).
- Lookahead and lookbehind
(?=…),(?!…),(?<=…),(?<!…), atomic groups(?>…), possessive quantifiers*+,++,?+ - Backreferences
\1,\k<name> - Java-only flags and escapes (
(?x),(?u),(?d),\h,\R,\Z,\e, and so on) - The
(?U)flag: in RE2 it swaps greedy and lazy repetition, but this engine cannot keep that meaning, so it returns 400. This differs from Prometheus.
How It Differs from Prometheus
How values are stored is different. Prometheus stores a counter as a monotonically increasing cumulative value.
APM stores each field in its own way. The engine looks at the measurement and the field together, finds how the field is stored, and calculates rate, irate, and increase for that storage.
The same field name can be stored differently in different measurements. For example, apdex_total is a delta, memory_total is a gauge, and jvm_classes_total is a cumulative value.
| Storage | Example metrics | rate(x[N]) | increase(x[N]) | irate(x[N]) | Warning |
|---|---|---|---|---|---|
| Already a rate (per-second value, time share in %) | tcp_activeOpens, interface_rxBytes, disk_diskReads, cpu_idle, WEB_stat_reqPerSec, MySQL qps* | avg_over_time(x[N]) | avg_over_time(x[N]) × N seconds | last_over_time(x[N]) | rate_alias |
| Delta (the increase since the previous collection) | apdex_count, gc_<collector>_gcCount, response status status*, MySQL diff* | Sum of the deltas in the window ÷ N seconds | Sum of the deltas in the window | The last delta ÷ the time between the last two samples | None |
| Cumulative (the raw counter value) | jvm_classes_total, memory_swapPageIn, Nginx accepts, HAProxy ereq, MySQL raw values such as comSelect | Same as Prometheus (reset correction, extrapolation to the window edges) | Same as Prometheus | Same as Prometheus (the last two samples) | None |
| Gauge, and fields that are not classified | jvm_heap_heapUsed, memory_used, OpenTelemetry metrics | Same as Prometheus | Same as Prometheus | Same as Prometheus | None |
tps,minTps, andmaxTps(such asapdex_tps) are per-second values, but APM stores them as integers and only for collection intervals that have transactions. The mean of the stored values is not the real throughput. So APM calculates them fromcount(a delta) of the same measurement.rate(apdex_tps[5m])gives the same value asrate(apdex_count[5m])and adds therate_aliaswarning. Query the throughput withrate(apdex_count[5m]).- A delta is the increase since the previous collection. So the sum of the deltas in the window is the increase in the window. APM does not extrapolate this sum. Extrapolation would drop the increase of the first sample in the window.
irateof a delta divides the last delta by the time between the last two samples. A metric such asapdex_counthas no samples in intervals without transactions, so the two samples can be far apart and the value is small.- Cumulative values, gauges, and fields that are not classified are calculated from the raw samples with the Prometheus formulas. The raw sample limit (
promql.raw.max.samples, default 1,000,000) applies.rateof a gauge has no meaning. Useavg_over_timeorderivon a gauge.
For a field already stored as a rate, use avg_over_time or last_over_time directly. A value with rate() and a value without it have the same unit but different values.
A query without a function gives the mean of the 5 minutes before the evaluation time, and rate(x[1m]) gives the mean of the last 1 minute. avg_over_time(x[1m]) gives the same value as rate(x[1m]) without a warning.
increase(x[N]) is the same as avg_over_time(x[N]) × N.
If the range N is shorter than the 2-second stored interval, a field already stored as a rate is read over 2 seconds, and APM uses one stored value as it is. The rate_short_window warning is then added.
A delta field keeps its value and only gets the rate_short_window warning. One delta is the increase over a period longer than the range, so the value can be larger.
The warnings are sentences in the warnings array of the response. Instant queries and range queries both have them, and Grafana shows them on the panel.
rate(tcp_activeOpens[5m]): tcp_activeOpens is stored as a rate (SYS agent per-second value), so the result is the mean of the stored rate over the window
rate(tcp_activeOpens[1s]): the range is shorter than the stored rate window 2s (the 2s stored interval); the stored rate values were used as they are
The calculation window is different. An instant query (/api/v1/query) and a range query (/api/v1/query_range) calculate values over these windows.
| Query | Instant query | Range query |
|---|---|---|
metric{...} without a function | The aggregate of the 5 minutes before the evaluation time (mean for a gauge, sum for a delta counter) | The same aggregate over one step from each time t, [t, t + step) |
rate(metric[5m]) | (evaluation time − 5m, evaluation time] calculated for the storage in the table above (a delta: sum ÷ 300) | (t − 5m, t] for each time t, calculated the same way |
increase(metric[5m]) | The same window calculated for the storage in the table above (a delta: sum) | (t − 5m, t] for each time t, calculated the same way |
avg_over_time(metric[5m]) and the others | Statistic over (evaluation time − 5m, evaluation time] | Statistic over (t − 5m, t] for each time t |
deriv(metric[5m]) and the other raw sample window functions | Calculated from the raw samples in (evaluation time − 5m, evaluation time] | Calculated from the raw samples in (t − 5m, t] for each time t |
max_over_time(<expr>[1h:1m]) and other subqueries | Calculated from the values of <expr> at the multiples of 1 minute in (evaluation time − 1h, evaluation time] | Calculated from the values of <expr> at the multiples of 1 minute in (t − 1h, t] for each time t |
In a range query, rate, irate, increase, and *_over_time calculate the (t − range, t] window at each time t, as Prometheus does, and put the value at time t.
The value is the same as an instant query at the same time. APM reads small windows, as long as the greatest common divisor of the range and the step, from InfluxDB and adds them up for each time.
In the following cases, APM calculates each step from one step window. The value at time t is the value of [t, t + step), and rate divides by the step seconds.
The response then has window function <function> approximated with step buckets (window <range>, step <step>) in warnings. Grafana shows this warning on the panel.
- One series would have more than 22,000 small windows,
(end − start + range) ÷ greatest common divisor. You can change the limit with the collection server system propertypromql.window.max.buckets. - The greatest common divisor of the range and the step is shorter than 1 second (for example,
[1500ms]with a 1-second step). - A calculated metric (
cpu_usage_percent= 100 − idle).
To avoid the warning, set a step that divides the range (step 1m or 5m for range 5m, and so on) or query a shorter period.
The raw sample window functions (delta, deriv, changes, quantile_over_time, and the others; see "Supported Functions" below) do not use small windows. They read each sample in the window.
rate, irate, and increase of cumulative values, gauges, and fields that are not classified, and irate of a delta field, read the samples the same way.
A range query reads (start − range, end] once and cuts it into the window of each time. These functions are never calculated from step windows, so they do not add warnings.
For a selector without a function, the range query value is the aggregate over one step from that time. This keeps the values the same as the existing APM screens. Prometheus uses the last sample before that time.
Supported:
- Label matchers
=,!=,=~,!~(a regular expression uses RE2 syntax and must match the full value) - Aggregation operators
sum,avg,min,max,count,stddev,stdvar,topk,bottomk,quantile,count_values,group, theby/withoutclauses, and nested aggregation - Scalar expressions as aggregation parameters (
topk(2 * 3, …),quantile(0.5 + 0.4, …),topk(scalar(x), …)) - The
offsetand@modifiers (metric offset 1h,rate(metric[5m] offset 1d),metric @ end()) - The label functions
label_replaceandlabel_join rate,irate,increase, and the*_over_timefunctions foravg,min,max,sum,count,last,present, andabsent- The raw sample window functions
delta,idelta,deriv,predict_linear,changes,resets,stddev_over_time,stdvar_over_time,quantile_over_time - Subqueries
<expr>[<range>:<resolution>](as the range argument of a window function; see "Subqueries" below) - Maths functions
abs,ceil,floor,round,clamp,clamp_min,clamp_max,sqrt,exp,ln,log2,log10,sgn - Trigonometric functions
sin,cos,tan,asin,acos,atan,sinh,cosh,tanh,asinh,acosh,atanh,deg,rad,pi() - Sorting with
sortandsort_desc, and checking that a series exists withabsent - Time and scalar conversion
time(),timestamp(),scalar(),vector(), and the date functionsminute,hour,day_of_week,day_of_month,day_of_year,days_in_month,month,year(KST by default, set withpromql.timezone) - Arithmetic operators (
+ - * / % ^ atan2), comparison operators (== != > < >= <=) with theboolmodifier, and unary minus - Vector matching
on(...)andignoring(...), and many-to-one / one-to-many matching withgroup_leftandgroup_right - The logical operators
and,or,unless - Hexadecimal numbers (
0x10= 16), and string literals in an instant query ("abc", result typestring;bad_datain a range query) - The HTTP API (
/api/v1/query,/api/v1/query_range,/api/v1/series,/api/v1/labels, and others)
Not supported -- a query that uses a query-syntax item gets a bad_data error (HTTP 400).
| Item | Note |
|---|---|
histogram_quantile() | APM does not store histogram buckets with an le label. For response times, query fields such as apdex_avgRT and apdex_maxRT |
mad_over_time(), holt_winters(), and other functions not in the list above | The error message shows the function name |
A subquery without a function (metric[1h:1m]), a subquery argument of rate, irate, or increase, and @ on a subquery | Use a subquery only as the range argument of a window function. APM calculates the rate family from how the stored field is stored, so it cannot use calculated values. The error message gives the reason |
__name__ regular expressions in a query, and selectors without a metric name ({__name__=~"a|b"}, {ip_addr="..."}) | Specify a metric name. {__name__="x"} is processed the same as x. /api/v1/series and match[] accept __name__ regular expressions (up to 20 matching metric names) |
Of the features that are not query syntax, recording and alerting rules are not supported. Automatic rollup for long-range queries is not applied. For queries over several weeks, specify a large step.
Note An operation between two different metrics (
a / b) gives an empty result. The default matching compares all labels except__name__, and the two metrics have differentnamespaceandmetric_namelabels. Specify the labels to compare, as injvm_heap_heapUsed / on(instance) cpu_usage_cpuUsage. Fields of the same metric (jvm_heap_heapUsed / jvm_heap_heapMax) do not need this.
Supported Functions
Instant vectors, range vectors, and scalars are supported. Use a range vector (metric[5m]) as the argument of a function such as rate.
Use string literals only as arguments of label_replace, label_join, and count_values, and as the result of an instant query.
Rate / counter functions -- range vector → instant vector:
rate(metric[5m]) # average rate of change per second (for each storage, see "How It Differs from Prometheus")
irate(metric[5m]) # instantaneous rate from the last two samples
increase(metric[5m]) # the total increase over the window
The over_time family -- window statistics:
avg_over_time(metric[1h]) min_over_time(metric[1h]) max_over_time(metric[1h])
sum_over_time(metric[1h]) count_over_time(metric[1h]) last_over_time(metric[1h])
present_over_time(metric[1h]) absent_over_time(metric[1h])
present_over_time is 1 for each series that has a value in the window. absent_over_time is 1 when no series has a value in the window.
Both functions are queried in the same way as count_over_time.
Raw sample window functions -- read each sample in the window and calculate with the same formulas as Prometheus:
delta(jvm_heap_heapUsed[10m]) # difference between the first and last sample values, extended to the window length
idelta(jvm_heap_heapUsed[5m]) # difference between the last two samples
deriv(jvm_heap_heapUsed[30m]) # change per second (slope of a least-squares regression)
predict_linear(jvm_heap_heapUsed[1h], 4 * 3600) # expected value 4 hours (14,400 seconds) later
changes(jvm_heap_heapUsed[1h]) # number of times two consecutive samples differ
resets(jvm_heap_heapUsed[1h]) # number of times the value goes down between two consecutive samples
stddev_over_time(jvm_heap_heapUsed[1h]) # standard deviation
stdvar_over_time(jvm_heap_heapUsed[1h]) # variance
quantile_over_time(0.95, jvm_heap_heapUsed[1h]) # 95th percentile of the window samples
delta,idelta,deriv, andpredict_linearhave a value only for a window with 2 or more samples; the others need 1 or more. A time without enough samples has no point in the result.deltaextends the difference to the window start when the first sample is within 1.1 times the average sample interval of it, and otherwise by half the average interval. The last sample and the window end work the same way.predict_linear(v[r], s)is the valuesseconds after the evaluation time.sand the φ ofquantile_over_timeare scalar expressions (4 * 3600,scalar(x)). In a range query, a value that differs per step is used at each step. A φ below 0 gives-Inf, and a φ above 1 gives+Inf.stddev_over_timeandstdvar_over_timeare the population standard deviation and variance of all samples in the window.- The result does not have the
__name__label. - One query can read up to 1,000,000 raw samples in total. Beyond that, the query gets a
bad_dataerror (HTTP 400)query reads too many samples (<read> > <limit>); use a shorter range or fewer series, and no approximate value. A metric collected every 2 seconds has 43,200 samples per series per day. You can change the limit with the collection server system propertypromql.raw.max.samples. - With counters: APM stores each counter field in its own way (see the storage table in "How It Differs from Prometheus" above).
delta,idelta,resets, andchangeson a metric stored as a delta (such asapdex_countorgc_<collector>_gcCount) calculate over the stored changes. Query the increase and the rate of a counter withincreaseandrate.
Subqueries -- <expr>[<range>:<resolution>] calculates <expr> at intervals of the resolution and uses the values as a range vector. Write it as the range argument of a window function:
max_over_time(sum(jvm_heap_heapUsed)[1h:1m]) # maximum of the heap sum in the last hour
avg_over_time((jvm_heap_heapUsed / jvm_heap_heapMax)[1h:1m]) # 1-hour average of the heap usage ratio
deriv(sum(jvm_heap_heapUsed)[30m:1m]) # change per second of the heap sum
max_over_time(rate(gc_G1_Young_Generation_gcCount[5m])[1d:5m] offset 1d) # maximum GC frequency during yesterday
- The functions that accept a subquery are the
*_over_timefunctions (avg,min,max,sum,count,last,present,absent) and the 9 raw sample window functions. - When you omit the resolution (
[1h:]), it is 1 minute. - APM calculates
<expr>at the times that are multiples of the resolution in unix time. These times do not depend on the query start, so a dashboard refresh does not change the value at the same time. - The value at time t comes from the calculated values in
(t − range, t]. Withoffset, the window is(t − offset − range, t − offset]. - A selector in
<expr>uses the last sample in the 5 minutes(τ − 5m, τ]before each calculation time τ (the same as Prometheus). When the resolution is shorter than the collection interval, the same sample value repeats. When there is no sample for more than 5 minutes, that time has no value. This is different from the range query value of a selector without a function (a step window aggregate). - A time without a value is not put in the window. It is not filled with 0.
- You can put a subquery in a subquery (
max_over_time(sum_over_time(x[30s:10s])[5m:10s])). - The result of
absent_over_time(<subquery>)has the labels{}(the same as Prometheus). The other functions use the same label rules as the usual window functions. - One subquery can calculate up to 50,000,000 points (series × calculation times). Beyond that, the query gets a
bad_dataerror (HTTP 400)subquery evaluates too many points (<points> > <limit>); use a coarser resolution. APM checks the number of times before the calculation and the number of points after it. You can change the limit with the collection server system propertypromql.subquery.max.points. - You cannot use a calculated metric (such as
cpu_usage_percent) in a subquery. - APM does not approximate the values in a subquery. In these cases, the query gets a
bad_dataerror (HTTP 400).- A selector in the subquery would read more than 22,000 small windows (
promql.window.max.buckets) for one series,(query period + range + 5m) ÷ (greatest common divisor of 5m and the resolution). The error message issubquery needs too many buckets per series (<windows> > <limit>). APM checks this before the calculation starts. Use a larger resolution or a shorter range. - The greatest common divisor of 5 minutes and the resolution is shorter than 1 second (for example,
[1m:700ms]). - A calculated metric is used in a subquery, or a window function in a subquery cannot be calculated exactly. The error message is
subquery cannot be evaluated exactly: ….
- A selector in the subquery would read more than 22,000 small windows (
Aggregation operators -- sum, avg, min, max, count, stddev, stdvar,
topk(N, …), bottomk(N, …), quantile(0.95, …), count_values("label", …), group.
Group with a by or without clause. A by result keeps only the listed labels. A without result removes the listed labels and __name__:
avg by (instance) (jvm_heap_heapUsed) # average per instance
sum without (user_key) (jvm_heap_heapUsed) # group by everything except user_key
topk(5, rate(gc_G1_Young_Generation_gcCount[5m])) # the top 5 instances by GC frequency
count(count by (ip_addr) (jvm_heap_heapUsed)) # number of hosts
count_values("heap_gib", round(jvm_heap_heapMax / 1024 / 1024 / 1024)) # number of instances for each maximum heap size (GiB)
The N of topk and bottomk and the value of quantile are scalar expressions (2 * 3, scalar(x)). In a range query, a value that differs per step is used at each step.
N is rounded down. If N is less than 1 or NaN, that step has no result.
A vector gives a bad_data error. Wrap it in scalar().
count_values writes the label value in the same format as Prometheus (1, 0.5, 1000000, NaN, +Inf).
Time shift -- offset and @:
jvm_heap_heapUsed offset 1h # the value one hour ago
rate(gc_G1_Young_Generation_gcCount[5m] offset 1d) # the GC frequency at the same time one day ago
jvm_heap_heapUsed @ 1700000000 # the value at a fixed time (unix seconds)
jvm_heap_heapUsed @ end() # the value at the end of the query range
Write offset directly after a selector, a range selector, or a subquery. After a function result or an expression in parentheses, offset gives a bad_data error.
A negative offset (offset -5m) moves forward in time.
@ reads the value at that one time. In a range query, all steps have the same value.
@ start() and @ end() are the start and end of a range query. In an instant query, both are the evaluation time.
You can use offset and @ together, in any order.
Label functions -- label_replace and label_join:
# the part of instance_id before "-" as the app label
label_replace(jvm_heap_heapUsed, "app", "$1", "instance_id", "(.*)-.*")
# ip_addr and instance_id joined with ":" as the host label
label_join(jvm_heap_heapUsed, "host", ":", "ip_addr", "instance_id")
The regular expression of label_replace uses the same RE2 syntax as label matchers and must match the full source label value. A series that does not match stays the same.
In the replacement, use $1, ${1}, $name, ${name}, and $$. Define a named group as (?P<name>…).
If the result is an empty string, the label is removed.
If the label_replace result has two or more series with the same labels, you get a bad_data error.
Maths, arithmetic, and comparison:
abs(metric) ceil(metric) floor(metric) round(metric) round(metric, 0.5)
clamp(metric, 0, 100) clamp_min(metric, 0) clamp_max(metric, 100)
sqrt(metric) exp(metric) ln(metric) log2(metric) log10(metric) sgn(metric)
sin(metric) cos(metric) tan(metric) asin(metric) acos(metric) atan(metric)
sinh(metric) cosh(metric) tanh(metric) asinh(metric) acosh(metric) atanh(metric)
deg(metric) rad(metric) pi()
jvm_heap_heapUsed / jvm_heap_heapMax
jvm_heap_heapUsed / on(instance) cpu_usage_cpuUsage
rate(metric[5m]) * 60
jvm_heap_heapUsed > 1024 * 1024 * 1024 # returns only the series over 1 GB
jvm_heap_heapUsed > bool 1024 * 1024 * 1024 # 1 (over) or 0 for each series
The result of an arithmetic operation, a maths function, rate, irate, increase, *_over_time, or a raw sample window function does not have the __name__ label (last_over_time keeps it).
In a range query, topk and bottomk rank the series again at every step. A series has only the points of the steps where it was picked, so a line in a graph can have gaps (the same as Prometheus). A comparison filter (without bool) keeps only the series that meet the condition, with their labels.
Division by zero gives the value +Inf, -Inf, or NaN.
A maths function given a value outside its domain keeps the series and gives NaN or -Inf (sqrt(-1) = NaN, ln(0) = -Inf).
pi() is the scalar π.
Sorting and checking for series -- sort, sort_desc, absent:
sort_desc(jvm_heap_heapUsed) # largest value first
absent(jvm_heap_heapUsed{instance_id="api-53-x"}) # {instance_id="api-53-x"} 1 when the series does not exist
absent_over_time(jvm_heap_heapUsed{instance_id="api-53-x"}[10m]) # 1 when there is no value for 10 minutes
sort and sort_desc order the result of an instant query by value. NaN comes last, and equal values are ordered by labels. A range query keeps its order (the same as Prometheus).
absent returns 1 at each time where no series of the argument has a value. When a value exists, the result is empty.
The result labels come from the = matchers of the argument selector (__name__ excluded, app_name included). A label with two or more matchers is left out, and an argument that is not a selector (absent(sum(…))) gives no labels.
Time and scalar conversion -- time, scalar, vector, date functions:
time() # evaluation time (unix seconds); the step time in a range query
time() - timestamp(jvm_heap_heapUsed) # seconds since the last collected sample
scalar(sum(jvm_heap_heapUsed)) # a vector with one series as a scalar
jvm_heap_heapUsed > scalar(avg(jvm_heap_heapUsed)) # series above the overall average at each step
topk(scalar(count(jvm_heap_heapUsed) / 7), jvm_heap_heapUsed) # top 1/7 of the instances
vector(time()) # a scalar as a series without labels
hour() # hour of the evaluation time (0-23, KST by default)
year(vector(1700000000)) # year of unix time 1700000000 = 2023
time()is the evaluation time in unix seconds. In a range query it is the time of each step.scalar(v)is, at each step, the value of the only series ofvwith a value there. With no series or two or more, it isNaN.vector(s)is one series without labels whose value at each step is the value ofs.- The date functions use KST (Asia/Seoul) by default:
minute(0-59),hour(0-23),day_of_week(0 = Sunday to 6),day_of_month(1-31),day_of_year(1-366),days_in_month(28-31),month(1-12),year.hour()is 9 at 9 a.m. KST. - Set the time zone with the JVM system property
promql.timezoneon the collection server. The value is a time zone name such asAsia/SeoulorUTC(for example-Dpromql.timezone=UTC). An invalid value is logged as a warning and KST is used. - Prometheus uses UTC, so with the default setting the result differs from Prometheus by 9 hours. Set
promql.timezone=UTCwhen you need the Prometheus value. observ (Cache PromQL) also defaults to KST. - Without an argument, a date function gives one series without labels computed from the evaluation time (the step time in a range query). With a vector argument, each value is read as unix seconds, and the result drops
__name__. Fractions of a second are dropped, andNaNand±Infstay as they are. - A scalar can have a different value at each step. Every place that takes a scalar (arithmetic and comparison operators,
round,clamp,clamp_min,clamp_max, the seconds ofpredict_linear, the φ ofquantile_over_time, the parameter oftopk,bottomk,quantile) uses the value of each step. A step where the minimum ofclampis greater than the maximum has no result. - When the query result is a scalar, an instant query returns result type
scalar, and a range query returns one series without labels. - When
vis a selector (also in parentheses),timestamp(v)is the time of the stored raw sample. At each time it returns the time of the last raw sample in the last 5 minutes ((t − offset − 5m, t − offset]) in unix seconds. Milliseconds are the decimal part. When there is no sample in 5 minutes, that time has no point (the same as Prometheus). timestamp(v)reads raw samples, so the raw sample limit applies (see "Raw sample window functions" above). A long range query over the limit gets abad_dataerror. With@, every time uses the same window. Inside a subquery the same rule applies at each inner calculation time.- When
vis not a selector (timestamp(sum(x)),timestamp(rate(x[5m]))), it is the time of each point, which is the evaluation time (the step time in a range query). The result drops__name__. A range vector (timestamp(x[5m])) gets abad_dataerror.
Logical operators -- and, or, unless:
jvm_heap_heapUsed and on(instance) jvm_heap_heapMax > 1024 * 1024 * 1024 # heap used of the instances whose maximum heap is over 1 GB
jvm_heap_heapUsed unless jvm_heap_heapUsed{instance_id="khan11"} # all except khan11
rate(gc_G1_Young_Generation_gcCount[5m]) > 1 or rate(gc_G1_Young_Generation_gcCount[5m]) < 0.1 # series that meet either condition
a and bkeeps only the points ofafor whichbhas a point with the same matching key at the same time. The values and labels are those ofa.a unless bkeeps only the points ofafor whichbhas no point with the same matching key at the same time.a or bgives all points ofa, plus the points ofbfor whichahas no point with the same matching key at the same time. Each keeps its own labels.
The matching key is the same as for arithmetic operators: all labels except __name__ by default, changed with on(...) or ignoring(...). The result keeps __name__.
A range query decides at each time separately. A scalar operand, or use with bool, group_left, or group_right, gives a bad_data error.
Many-to-one matching -- group_left, group_right:
jvm_heap_heapUsed / on(ip_addr) group_left cpu_usage_cpuUsage # several instances on one host
jvm_heap_heapUsed / on(instance) group_left(app_name) jvm_heap_heapMax # copies the app_name label of the right side to the result
Use group_left when the left side has several series with the same matching key and the right side has one. group_right is the reverse.
It comes after on(...) or ignoring(...), and works only with arithmetic and comparison operators.
- The result labels are the labels of the side with several series (the left side for
group_left). Unlike one-to-one matching, they are not reduced to theon(...)labels. Arithmetic operators andboolcomparisons remove__name__. - Each label listed in the parentheses is copied from the side with one series. When that side does not have the label, it is removed from the result.
- When the side that must have one series has two or more series with the same matching key at the same time, you get a
bad_dataerror (many-to-many matching not allowed). Two or more result series with the same labels also give an error. - The value of a comparison filter (without
bool) is the value of the left operand (the same as Prometheus). - To put the right operand in parentheses, write an empty label list first, as in
group_left() (a + b). Prometheus reads parentheses right aftergroup_leftas the label list.
Operator precedence -- an operator higher in the table binds first (the same as Prometheus).
| Order | Operators | Grouping |
|---|---|---|
| 1 | ^ | Right to left (2 ^ 3 ^ 2 = 2 ^ (3 ^ 2) = 512) |
| 2 | Unary +, - | (-2 ^ 2 = -(2 ^ 2) = -4) |
| 3 | *, /, %, atan2 | Left to right |
| 4 | +, - | Left to right |
| 5 | ==, !=, <=, <, >=, > | Left to right |
| 6 | and, unless | Left to right |
| 7 | or | Left to right |
Versions before 2026-10-05 calculated -2 ^ 2 as 4 and 2 ^ 3 ^ 2 as 64. Dashboards that use such expressions show different values.
a atan2 b is the arc tangent of a over b for each point (radians, −π to π). As with the arithmetic operators, you can use on(...), ignoring(...), group_left, and group_right, and the result drops __name__. With bool it is a syntax error (bad_data).
Metric Explorer
A UI for running PromQL interactively in a browser -- reach it at the address in How to open it above.
| Area | Function |
|---|---|
| Metric selection dropdown | Automatic discovery of registered metrics (by __name__) |
| PromQL query input | Type it directly -- choosing from the dropdown fills it in automatically |
| Label filter panel | Dropdowns per ip_addr, agent_type, instance_id, and namespace |
| Aggregate function selection | sum, avg, topk, quantile, and others |
| Time range | Quick buttons from the last 5m to 7d, plus a custom range |
| Step | Auto / 15s / 1m / 5m / 10m / 30m / 1h |
| Result visualization | Line, area, and bar charts plus a table of raw values |
| Copy API URL | Copies the current query as an /api/v1/query_range URL -- for external integration such as Grafana |
Making your first chart in 30 seconds:
- Select
jvm_heap_heapUsedin the metric dropdown - Select
agent_type=WASin the label filter - Click the 1h time range and leave Step at Auto
- Click Run → a graph of one hour of heap usage across the WAS instances appears
Note Auto is usually enough for Step. A small step with a long range loads the server heavily.
Grafana integration -- the same API can be connected as a Prometheus data source in Grafana:
URL: http://{apm-server}/monitoring
Access: Server (default)
Grafana calls /api/v1/query, /api/v1/query_range, /api/v1/labels, and the rest automatically --
reuse existing dashboard and alerting assets as they are.
HTTP API Summary
| Method | Path | Purpose |
|---|---|---|
| GET / POST | /api/v1/query | Instant (single point in time) query |
| GET / POST | /api/v1/query_range | Range (time range) query |
| GET | /api/v1/series | Series metadata |
| GET | /api/v1/labels | The list of label names |
| GET | /api/v1/label/{name}/values | The list of values for a particular label |
| GET | /api/v1/metadata | Metric metadata |
| GET | /api/v1/metric-explorer | The Metric Explorer UI |
Parameters -- query (required), time for instant, and start, end, and step for range
(calculated automatically when omitted). /series, /labels, and /label/{name}/values accept match[], start, end, and limit.
With match[], only the label names and values of the series that match the selector are returned (the Grafana variable label_values(metric, label) uses this call).
limit is the number of entries kept from the start of the sorted result (0 or omitted means no limit; a negative value is bad_data). Exploratory calls:
# The full list of available metrics (~250)
curl "http://{apm-server}/monitoring/api/v1/label/__name__/values"
# The list of label names / the values of a particular label
curl "http://{apm-server}/monitoring/api/v1/labels"
curl "http://{apm-server}/monitoring/api/v1/label/agent_type/values"
# The labels of one metric / 10 instance values of that metric
curl -G "http://{apm-server}/monitoring/api/v1/labels" --data-urlencode "match[]=jvm_heap_heapUsed"
curl -G "http://{apm-server}/monitoring/api/v1/label/instance/values" \
--data-urlencode "match[]=jvm_heap_heapUsed" --data-urlencode "limit=10"
When match[] has several selectors and the storage query of one of them fails, the whole request is an error, with no partial result (the same as Prometheus).
HTTP error status -- the errorType and HTTP status of an error response are the same as in Prometheus, except unavailable. All APIs use the same rule.
errorType | HTTP | Cause |
|---|---|---|
bad_data | 400 | Syntax error, unsupported function, query size limit exceeded, incorrect parameter |
execution | 422 | Other execution errors (internal error of the collection server) |
internal | 500 | The storage (InfluxDB) answered with an error. The message starts with InfluxDB query failed: … |
timeout | 503 | The storage did not answer within the time limit, or the turn of a heavy query did not come within 10 seconds |
unavailable | 503 | The collection server could not connect to the storage. Prometheus uses 500 for this errorType. APM uses 503 because you can try again later |
For a 503, try again a little later.
A Collection of Frequently Used Queries
The metric names are examples following the general pattern -- check the exact names for your environment in the Metric Explorer.
JVM / WAS:
jvm_heap_heapUsed{agent_type="WAS"} # current heap usage across all WAS
avg_over_time(jvm_heap_heapUsed[5m]) # 5-minute average heap usage
topk(5, jvm_heap_heapUsed) # the top 5 instances by heap usage
(jvm_heap_heapUsed / jvm_heap_heapMax) > 0.8 # heap utilization over 80%
sum_over_time(gc_G1_Young_Generation_gcTime[1h]) # total GC time (1 hour)
rate(gc_G1_Young_Generation_gcCount[5m]) # GC count per second
datasource_pool_active{agent_type="WAS"} # DataSource connections in use
System:
topk(10, cpu_usage_cpuUsage{agent_type="SYS"}) # the top 10 hosts by CPU
cpu_usage_systemCpuUsage # system-wide CPU
avg_over_time(cpu_usage_cpuUsage[1h]) # 1-hour average CPU per instance
Aggregating per host:
avg by (ip_addr) (jvm_heap_heapUsed) # average heap per host
count by (agent_type) (jvm_heap_heapUsed) # instance count per agent_type
jvm_heap_heapUsed{ip_addr="192.168.80.190"} # heap usage for a particular IP
When It Does Not Work
| Symptom | What to check |
|---|---|
| You do not know the metric name | Check the actual name in the Metric Explorer dropdown or at /api/v1/label/__name__/values |
| The result is empty | A typo in a label value -- check the possible values at /api/v1/label/{name}/values, and check the case of agent_type |
| Metric Explorer sign-in fails | An API access key is required -- create and copy the key under Settings ▸ Users |
A rate() value is not what you expect | The calculation depends on how the field is stored -- see the storage table in "How It Differs from Prometheus" above. For a field already stored as a rate, a query without a function (the mean of the last 5 minutes) and rate(x[N]) (the mean of the last N) give different values |
The response warnings has … is stored as a rate … (rate_alias) | The field is already stored as a rate, so rate was calculated as avg_over_time and irate as last_over_time -- use avg_over_time or last_over_time directly to get no warning. The tps fields are calculated from count |
The response warnings has the range is shorter than … (rate_short_window) | The range is shorter than the 2-second stored interval -- set a range of 2 seconds or more |
| The sum or average does not match | A query without a function gives the mean or sum of the last 5 minutes (instant query) or of one step from each time (range query) -- write the window you want with sum_over_time or avg_over_time |
The response warnings has window function … approximated with step buckets | A window function was calculated from one step window -- set a step that divides the range or query a shorter period (see "The calculation window" above) |
query selects too many series error | One selector or window function reads more series than the limit (2,000,000 ÷ rows per series) -- use a larger step, shorten the query period, or narrow the series with labels (see "Querying a Specified Period" above) |
query reads too many rows error | The rows one query reads in total are more than the limit (8,000,000 by default) -- use a larger step, shorten the query period, or narrow the series with labels |
too many concurrent heavy queries; retry later error (HTTP 503) | 4 queries with more than 5,000 rows per series are running, and the turn did not come within 10 seconds -- run the query again a little later, or specify a larger step |
internal error (HTTP 500) InfluxDB query failed: … | The storage answered with an error -- check the InfluxDB error text in the message together with the InfluxQL in the collection server log |
timeout or unavailable error (HTTP 503) | The storage did not answer within the time limit, or it could not be reached -- check the InfluxDB status and run the query again a little later. If timeouts continue, shorten the range or specify a larger step |
query reads too many samples error | A calculation that reads raw samples (a raw sample window function, the rate family calculated from raw samples, timestamp()) reads more samples than the limit (1,000,000 by default) -- shorten the range or the query period, or narrow the series with labels |
subquery evaluates too many points error | A subquery calculates more points than the limit (50,000,000 by default) -- use a larger resolution, or shorten the query period or the range |
subquery needs too many buckets per series error | A selector in the subquery reads more small windows than the limit per series (22,000 by default) -- use a larger resolution, or shorten the range or the query period |
cannot be evaluated exactly error (subquery … cannot be evaluated exactly) | APM cannot calculate the values in the subquery exactly at the resolution -- use a larger resolution or a shorter range. A calculated metric cannot be used in a subquery (see "Subqueries" above) |
| An operation between two metrics is empty | The labels are different, so the series do not match -- specify the labels to compare, as in on(instance) |
bad_data error (HTTP 400) | Unsupported syntax or function -- check the function name in the error message and the "Not supported" table above |
histogram_quantile does not work | Not supported -- APM does not store le buckets. Query response-time fields such as apdex_avgRT and apdex_maxRT directly |
| A long-range query is slow | Specify a large step (30m to 1h) -- a small step over a long range is heavy |
Related Documents
- R2. Chart Metric Reference -- what each metric means, the signals of trouble, and thresholds
- E3. What the Metrics Mean -- the concepts of APDEX, response time, and throughput
- H3. My Dashboard -- gathering charts inside the console (no PromQL needed)
- H18. Managing Users, Groups, and Permissions -- where the API access key is issued
This chapter also appears with the same content as chapter 7 of the OPENMARU APM API Integration Guide. To read it alongside the JSON APIs that use a session cookie (graph data, configuration, user management), see that guide.