Skip to content

7. PromQL Integration

If you are already comfortable with Grafana dashboards or PromQL, the WAS, JVM, system, WEB, and DBMS metrics OPENMARU APM collects can be queried directly with standard PromQL-compatible queries. Collection, storage, querying, and visualization all happen in APM alone, with no separate Prometheus server, exporter, or scrape configuration, and existing Grafana dashboard assets can be reused.

  • APM's fine-grained metrics as they are -- reach JVM, GC, thread, DB connection pool, and transaction metrics with standard PromQL
  • Rich labels -- filter and aggregate precisely with APM context such as instance_id, ip_addr, and agent_type
  • rate/increase follow how a metric is stored -- a metric stored as a delta is summed over the window, and a metric already stored as a per-second value is averaged over the window
  • Relative time expressions -- intuitive forms such as now and now-1h are supported (a convenience standard Prometheus does not have)

For what each metric means, see OPENMARU APM User Guide R2. Chart Metric Reference and OPENMARU APM User Guide E3. What the Metrics Mean. To gather charts inside the console, OPENMARU APM User Guide H3. My Dashboard is enough -- PromQL is the tool for external integration and automation.

How to open it -- in a browser, http://{apm-server}/monitoring/api/v1/metric-explorer (the Metric Explorer UI). Sign in with an API access key -- go to the left menu ▸ Settings ▸ Users, press the edit user button, and create and copy the key.


Getting Started in 5 Minutes​

The simplest query -- one line gives the JVM heap usage (current value) of every WAS instance.

jvm_heap_heapUsed

To see only a particular instance, or only a particular IP range (regular expression):

jvm_heap_heapUsed{instance_id="apm-was-01"}
jvm_heap_heapUsed{ip_addr=~"192\\.168\\.80\\..*"}

Apply time functions to get the maximum over the last hour and the 5-minute average GC frequency:

max_over_time(jvm_heap_heapUsed[1h])
rate(gc_G1_Young_Generation_gcCount[5m])

Note Choose the exact metric name from the Metric Explorer dropdown, or check it with GET /api/v1/label/__name__/values (about 250 of them). GC metrics are named gc_<collector>_gcCount and gc_<collector>_gcTime; the collector name depends on the JVM GC setting (for example gc_G1_Young_Generation_gcCount, gc_PS_Scavenge_gcCount). The examples in this document use G1.


Querying a Specified Period​

A range query (/api/v1/query_range) takes the start and end times and a step (the data interval). Four time formats are supported.

FormatExampleNote
Relative timenow, now-1h, now-30m, now-1dThe most convenient -- units s/m/h/d
Unix epoch (seconds)1704067200shell date +%s
Unix epoch with a fraction1704067200.123Millisecond precision
RFC33392026-05-10T09:00:00+09:00Can include a time zone

Frequently used call examples:

# The last hour -- the most common case (step automatically 15s)
curl -G "http://{apm-server}/monitoring/api/v1/query_range" \
--data-urlencode "query=jvm_heap_heapUsed" \
--data-urlencode "start=now-1h" \
--data-urlencode "end=now"

# The weekly trend over the last 7 days (step automatically 5m)
curl -G "http://{apm-server}/monitoring/api/v1/query_range" \
--data-urlencode "query=max_over_time(cpu_usage_cpuUsage[1h])" \
--data-urlencode "start=now-7d" \
--data-urlencode "end=now"

# Specifying the step directly (15-second interval)
curl -G "http://{apm-server}/monitoring/api/v1/query_range" \
--data-urlencode "query=rate(gc_G1_Young_Generation_gcCount[5m])" \
--data-urlencode "start=now-30m" \
--data-urlencode "end=now" \
--data-urlencode "step=15s"

# A particular date range (KST)
curl -G "http://{apm-server}/monitoring/api/v1/query_range" \
--data-urlencode "query=jvm_heap_heapUsed{agent_type=\"WAS\"}" \
--data-urlencode "start=2026-05-10T09:00:00+09:00" \
--data-urlencode "end=2026-05-11T18:00:00+09:00"

Without a step, it is calculated automatically from the query range:

Query rangeAutomatic stepData points (example)Suggested use
12 hours or less15s1 hour ≈ 240Real time · standard dashboards
12 hours to 2 days1m24 hours ≈ 1,440Day-to-day comparison
2 to 8 days5m7 days ≈ 2,016Weekly trend
8 to 21 days10mFortnightly trend
21 days or more30m30 days ≈ 1,440Monthly comparison

Note The raw data is stored every 2 seconds. A smaller step is denser but loads the server more -- start with a large step and reduce it when you need to.

Range queries follow these rules.

  • The timestamps in the response are aligned to step boundaries. When start is 09:00:30 and the step is 1m, the first point is 09:00:00. This is because the storage groups values by step boundaries.
  • If end is before start, you get a bad_data error (HTTP 400).
  • If (end − start) ÷ step is more than 11,000, you get a bad_data error (HTTP 400). Specify a larger step. When the step is omitted, it is increased so that it stays within this limit (ranges longer than about 229 days).

The amount one query reads from the storage is also limited. The same limits apply to instant queries.

  • One selector or window function can read up to 2,000,000 rows (series × rows per series). The rows per series are the number of steps for a selector without a function, the number of small windows for a window function, and 1 for an instant query. So the number of series it can read is 2,000,000 ÷ rows per series. For example, a 3-day range with a 30-second step has 8,641 rows per series, so up to 231 series. Beyond that, the query gets a bad_data error (HTTP 400) query selects too many series (more than <series limit> for this range and step); use a coarser step, a shorter range or a narrower selector. The storage is asked for at most one series more than the limit, so most of the extra series are rejected before they are read. This request count applies per storage unit (measurement), so a metric split across several measurements can read that many more. Use a larger step, shorten the query period, or narrow the series with labels. You can change the limit with the collection server system property promql.leaf.max.rows. Calculations that read raw samples (the raw sample window functions, the rate, irate, and increase calculated from raw samples in "How It Differs from Prometheus" below, and timestamp()) use the raw sample limit instead of the limits in this section (see "Supported Functions" below).
  • All selectors and window functions of one query can read up to 8,000,000 rows in total. Beyond that, the query gets a bad_data error (HTTP 400) query reads too many rows (<rows read> > <limit>); use a coarser step, a shorter range or a narrower selector. You can change the limit with the collection server system property promql.query.max.rows.
  • Selectors and window functions with more than 5,000 rows per series (for example, a 3-day range with a 30-second step) read the storage at most 4 at a time across the whole collection server. When 4 are already reading, the next one waits up to 10 seconds. If its turn still does not come, the query gets a timeout error (HTTP 503) too many concurrent heavy queries; retry later. Run the query again a little later. Queries with 5,000 rows per series or fewer (such as the usual dashboard panels) do not wait. You can change the threshold, the number at a time, and the wait time with the collection server system properties promql.heavy.min.rows.per.series, promql.heavy.max.concurrent, and promql.heavy.wait.ms (milliseconds). promql.heavy.max.concurrent is read once when the collection server starts.

If you need only one value at a single point in time, use an instant query (/api/v1/query):

# The heap usage now / one hour ago
curl -G "http://{apm-server}/monitoring/api/v1/query" \
--data-urlencode "query=jvm_heap_heapUsed"
curl -G "http://{apm-server}/monitoring/api/v1/query" \
--data-urlencode "query=jvm_heap_heapUsed" \
--data-urlencode "time=now-1h"

Metric Names and Labels​

APM metrics are distinguished by five label dimensions. They are used directly in query filters and aggregation.

LabelMeaningExample values
__name__The metric name -- in the form {nameSpace}_{metricName}_{field}jvm_heap_heapUsed, cpu_usage_cpuUsage
ip_addrThe IP of the host the agent is installed on192.168.80.190
agent_typeThe agent typeWAS, SYS, DBMS, WEB
instance_idThe instance identifier (unique within the host)apm-was-01, nginx-1
namespaceThe metric groupjvm, transaction, cpu, memory

Prometheus-standard labels are generated automatically as well -- instance (= {ip_addr}:{instance_id}) and job (= agent_type).

The app_name label is the application name of the instance. A WAS instance gets the application name its agent reports; the name of a custom group created by a user is never used. To filter by a custom group, use a matcher such as metric{app_name="<group name>"}.

For metric names, only one form has to be known: {nameSpace}_{metricName}_{field} (the trailing _field may be omitted -- omitting it selects the default field).

ExampleMeaning
jvm_heap_heapUsedThe used field of the JVM heap
jvm_heap_heapMaxThe maximum field of the same metric
cpu_usage_cpuUsageCPU utilization
cpu_usage_systemCpuUsageSystem-wide CPU utilization

All four standard label matchers are supported, and several conditions are combined with commas:

metric{label="value"} # equal
metric{label!="value"} # not equal
metric{label=~"regex"} # regular expression match
metric{label!~"regex"} # regular expression non-match

jvm_heap_heapUsed{ip_addr="192.168.80.190", instance_id=~"apm-was-.*", agent_type="WAS"}

Regular expressions use the same RE2 syntax as Prometheus (Go regexp) and must match the full value. POSIX character classes such as [[:alpha:]], named groups (?P<name>…), and the flags (?i), (?m), (?s) are accepted. The following syntax, which RE2 does not accept, returns a 400 error (invalid regular expression …: <reason>).

  • Lookahead and lookbehind (?=…), (?!…), (?<=…), (?<!…), atomic groups (?>…), possessive quantifiers *+, ++, ?+
  • Backreferences \1, \k<name>
  • Java-only flags and escapes ((?x), (?u), (?d), \h, \R, \Z, \e, and so on)
  • The (?U) flag: in RE2 it swaps greedy and lazy repetition, but this engine cannot keep that meaning, so it returns 400. This differs from Prometheus.

How It Differs from Prometheus​

How values are stored is different. Prometheus stores a counter as a monotonically increasing cumulative value. APM stores each field in its own way. The engine looks at the measurement and the field together, finds how the field is stored, and calculates rate, irate, and increase for that storage. The same field name can be stored differently in different measurements. For example, apdex_total is a delta, memory_total is a gauge, and jvm_classes_total is a cumulative value.

StorageExample metricsrate(x[N])increase(x[N])irate(x[N])Warning
Already a rate (per-second value, time share in %)tcp_activeOpens, interface_rxBytes, disk_diskReads, cpu_idle, WEB_stat_reqPerSec, MySQL qps*avg_over_time(x[N])avg_over_time(x[N]) × N secondslast_over_time(x[N])rate_alias
Delta (the increase since the previous collection)apdex_count, gc_<collector>_gcCount, response status status*, MySQL diff*Sum of the deltas in the window ÷ N secondsSum of the deltas in the windowThe last delta ÷ the time between the last two samplesNone
Cumulative (the raw counter value)jvm_classes_total, memory_swapPageIn, Nginx accepts, HAProxy ereq, MySQL raw values such as comSelectSame as Prometheus (reset correction, extrapolation to the window edges)Same as PrometheusSame as Prometheus (the last two samples)None
Gauge, and fields that are not classifiedjvm_heap_heapUsed, memory_used, OpenTelemetry metricsSame as PrometheusSame as PrometheusSame as PrometheusNone
  • tps, minTps, and maxTps (such as apdex_tps) are per-second values, but APM stores them as integers and only for collection intervals that have transactions. The mean of the stored values is not the real throughput. So APM calculates them from count (a delta) of the same measurement. rate(apdex_tps[5m]) gives the same value as rate(apdex_count[5m]) and adds the rate_alias warning. Query the throughput with rate(apdex_count[5m]).
  • A delta is the increase since the previous collection. So the sum of the deltas in the window is the increase in the window. APM does not extrapolate this sum. Extrapolation would drop the increase of the first sample in the window.
  • irate of a delta divides the last delta by the time between the last two samples. A metric such as apdex_count has no samples in intervals without transactions, so the two samples can be far apart and the value is small.
  • Cumulative values, gauges, and fields that are not classified are calculated from the raw samples with the Prometheus formulas. The raw sample limit (promql.raw.max.samples, default 1,000,000) applies. rate of a gauge has no meaning. Use avg_over_time or deriv on a gauge.

For a field already stored as a rate, use avg_over_time or last_over_time directly. A value with rate() and a value without it have the same unit but different values. A query without a function gives the mean of the 5 minutes before the evaluation time, and rate(x[1m]) gives the mean of the last 1 minute. avg_over_time(x[1m]) gives the same value as rate(x[1m]) without a warning. increase(x[N]) is the same as avg_over_time(x[N]) × N.

If the range N is shorter than the 2-second stored interval, a field already stored as a rate is read over 2 seconds, and APM uses one stored value as it is. The rate_short_window warning is then added. A delta field keeps its value and only gets the rate_short_window warning. One delta is the increase over a period longer than the range, so the value can be larger.

The warnings are sentences in the warnings array of the response. Instant queries and range queries both have them, and Grafana shows them on the panel.

rate(tcp_activeOpens[5m]): tcp_activeOpens is stored as a rate (SYS agent per-second value), so the result is the mean of the stored rate over the window
rate(tcp_activeOpens[1s]): the range is shorter than the stored rate window 2s (the 2s stored interval); the stored rate values were used as they are

The calculation window is different. An instant query (/api/v1/query) and a range query (/api/v1/query_range) calculate values over these windows.

QueryInstant queryRange query
metric{...} without a functionThe aggregate of the 5 minutes before the evaluation time (mean for a gauge, sum for a delta counter)The same aggregate over one step from each time t, [t, t + step)
rate(metric[5m])(evaluation time − 5m, evaluation time] calculated for the storage in the table above (a delta: sum ÷ 300)(t − 5m, t] for each time t, calculated the same way
increase(metric[5m])The same window calculated for the storage in the table above (a delta: sum)(t − 5m, t] for each time t, calculated the same way
avg_over_time(metric[5m]) and the othersStatistic over (evaluation time − 5m, evaluation time]Statistic over (t − 5m, t] for each time t
deriv(metric[5m]) and the other raw sample window functionsCalculated from the raw samples in (evaluation time − 5m, evaluation time]Calculated from the raw samples in (t − 5m, t] for each time t
max_over_time(<expr>[1h:1m]) and other subqueriesCalculated from the values of <expr> at the multiples of 1 minute in (evaluation time − 1h, evaluation time]Calculated from the values of <expr> at the multiples of 1 minute in (t − 1h, t] for each time t

In a range query, rate, irate, increase, and *_over_time calculate the (t − range, t] window at each time t, as Prometheus does, and put the value at time t. The value is the same as an instant query at the same time. APM reads small windows, as long as the greatest common divisor of the range and the step, from InfluxDB and adds them up for each time.

In the following cases, APM calculates each step from one step window. The value at time t is the value of [t, t + step), and rate divides by the step seconds. The response then has window function <function> approximated with step buckets (window <range>, step <step>) in warnings. Grafana shows this warning on the panel.

  • One series would have more than 22,000 small windows, (end − start + range) ÷ greatest common divisor. You can change the limit with the collection server system property promql.window.max.buckets.
  • The greatest common divisor of the range and the step is shorter than 1 second (for example, [1500ms] with a 1-second step).
  • A calculated metric (cpu_usage_percent = 100 − idle).

To avoid the warning, set a step that divides the range (step 1m or 5m for range 5m, and so on) or query a shorter period.

The raw sample window functions (delta, deriv, changes, quantile_over_time, and the others; see "Supported Functions" below) do not use small windows. They read each sample in the window. rate, irate, and increase of cumulative values, gauges, and fields that are not classified, and irate of a delta field, read the samples the same way. A range query reads (start − range, end] once and cuts it into the window of each time. These functions are never calculated from step windows, so they do not add warnings.

For a selector without a function, the range query value is the aggregate over one step from that time. This keeps the values the same as the existing APM screens. Prometheus uses the last sample before that time.

Supported:

  • Label matchers =, !=, =~, !~ (a regular expression uses RE2 syntax and must match the full value)
  • Aggregation operators sum, avg, min, max, count, stddev, stdvar, topk, bottomk, quantile, count_values, group, the by / without clauses, and nested aggregation
  • Scalar expressions as aggregation parameters (topk(2 * 3, …), quantile(0.5 + 0.4, …), topk(scalar(x), …))
  • The offset and @ modifiers (metric offset 1h, rate(metric[5m] offset 1d), metric @ end())
  • The label functions label_replace and label_join
  • rate, irate, increase, and the *_over_time functions for avg, min, max, sum, count, last, present, and absent
  • The raw sample window functions delta, idelta, deriv, predict_linear, changes, resets, stddev_over_time, stdvar_over_time, quantile_over_time
  • Subqueries <expr>[<range>:<resolution>] (as the range argument of a window function; see "Subqueries" below)
  • Maths functions abs, ceil, floor, round, clamp, clamp_min, clamp_max, sqrt, exp, ln, log2, log10, sgn
  • Trigonometric functions sin, cos, tan, asin, acos, atan, sinh, cosh, tanh, asinh, acosh, atanh, deg, rad, pi()
  • Sorting with sort and sort_desc, and checking that a series exists with absent
  • Time and scalar conversion time(), timestamp(), scalar(), vector(), and the date functions minute, hour, day_of_week, day_of_month, day_of_year, days_in_month, month, year (KST by default, set with promql.timezone)
  • Arithmetic operators (+ - * / % ^ atan2), comparison operators (== != > < >= <=) with the bool modifier, and unary minus
  • Vector matching on(...) and ignoring(...), and many-to-one / one-to-many matching with group_left and group_right
  • The logical operators and, or, unless
  • Hexadecimal numbers (0x10 = 16), and string literals in an instant query ("abc", result type string; bad_data in a range query)
  • The HTTP API (/api/v1/query, /api/v1/query_range, /api/v1/series, /api/v1/labels, and others)

Not supported -- a query that uses a query-syntax item gets a bad_data error (HTTP 400).

ItemNote
histogram_quantile()APM does not store histogram buckets with an le label. For response times, query fields such as apdex_avgRT and apdex_maxRT
mad_over_time(), holt_winters(), and other functions not in the list aboveThe error message shows the function name
A subquery without a function (metric[1h:1m]), a subquery argument of rate, irate, or increase, and @ on a subqueryUse a subquery only as the range argument of a window function. APM calculates the rate family from how the stored field is stored, so it cannot use calculated values. The error message gives the reason
__name__ regular expressions in a query, and selectors without a metric name ({__name__=~"a|b"}, {ip_addr="..."})Specify a metric name. {__name__="x"} is processed the same as x. /api/v1/series and match[] accept __name__ regular expressions (up to 20 matching metric names)

Of the features that are not query syntax, recording and alerting rules are not supported. Automatic rollup for long-range queries is not applied. For queries over several weeks, specify a large step.

Note An operation between two different metrics (a / b) gives an empty result. The default matching compares all labels except __name__, and the two metrics have different namespace and metric_name labels. Specify the labels to compare, as in jvm_heap_heapUsed / on(instance) cpu_usage_cpuUsage. Fields of the same metric (jvm_heap_heapUsed / jvm_heap_heapMax) do not need this.


Supported Functions​

Instant vectors, range vectors, and scalars are supported. Use a range vector (metric[5m]) as the argument of a function such as rate. Use string literals only as arguments of label_replace, label_join, and count_values, and as the result of an instant query.

Rate / counter functions -- range vector → instant vector:

rate(metric[5m]) # average rate of change per second (for each storage, see "How It Differs from Prometheus")
irate(metric[5m]) # instantaneous rate from the last two samples
increase(metric[5m]) # the total increase over the window

The over_time family -- window statistics:

avg_over_time(metric[1h]) min_over_time(metric[1h]) max_over_time(metric[1h])
sum_over_time(metric[1h]) count_over_time(metric[1h]) last_over_time(metric[1h])
present_over_time(metric[1h]) absent_over_time(metric[1h])

present_over_time is 1 for each series that has a value in the window. absent_over_time is 1 when no series has a value in the window. Both functions are queried in the same way as count_over_time.

Raw sample window functions -- read each sample in the window and calculate with the same formulas as Prometheus:

delta(jvm_heap_heapUsed[10m]) # difference between the first and last sample values, extended to the window length
idelta(jvm_heap_heapUsed[5m]) # difference between the last two samples
deriv(jvm_heap_heapUsed[30m]) # change per second (slope of a least-squares regression)
predict_linear(jvm_heap_heapUsed[1h], 4 * 3600) # expected value 4 hours (14,400 seconds) later
changes(jvm_heap_heapUsed[1h]) # number of times two consecutive samples differ
resets(jvm_heap_heapUsed[1h]) # number of times the value goes down between two consecutive samples
stddev_over_time(jvm_heap_heapUsed[1h]) # standard deviation
stdvar_over_time(jvm_heap_heapUsed[1h]) # variance
quantile_over_time(0.95, jvm_heap_heapUsed[1h]) # 95th percentile of the window samples
  • delta, idelta, deriv, and predict_linear have a value only for a window with 2 or more samples; the others need 1 or more. A time without enough samples has no point in the result.
  • delta extends the difference to the window start when the first sample is within 1.1 times the average sample interval of it, and otherwise by half the average interval. The last sample and the window end work the same way.
  • predict_linear(v[r], s) is the value s seconds after the evaluation time. s and the φ of quantile_over_time are scalar expressions (4 * 3600, scalar(x)). In a range query, a value that differs per step is used at each step. A φ below 0 gives -Inf, and a φ above 1 gives +Inf.
  • stddev_over_time and stdvar_over_time are the population standard deviation and variance of all samples in the window.
  • The result does not have the __name__ label.
  • One query can read up to 1,000,000 raw samples in total. Beyond that, the query gets a bad_data error (HTTP 400) query reads too many samples (<read> > <limit>); use a shorter range or fewer series, and no approximate value. A metric collected every 2 seconds has 43,200 samples per series per day. You can change the limit with the collection server system property promql.raw.max.samples.
  • With counters: APM stores each counter field in its own way (see the storage table in "How It Differs from Prometheus" above). delta, idelta, resets, and changes on a metric stored as a delta (such as apdex_count or gc_<collector>_gcCount) calculate over the stored changes. Query the increase and the rate of a counter with increase and rate.

Subqueries -- <expr>[<range>:<resolution>] calculates <expr> at intervals of the resolution and uses the values as a range vector. Write it as the range argument of a window function:

max_over_time(sum(jvm_heap_heapUsed)[1h:1m]) # maximum of the heap sum in the last hour
avg_over_time((jvm_heap_heapUsed / jvm_heap_heapMax)[1h:1m]) # 1-hour average of the heap usage ratio
deriv(sum(jvm_heap_heapUsed)[30m:1m]) # change per second of the heap sum
max_over_time(rate(gc_G1_Young_Generation_gcCount[5m])[1d:5m] offset 1d) # maximum GC frequency during yesterday
  • The functions that accept a subquery are the *_over_time functions (avg, min, max, sum, count, last, present, absent) and the 9 raw sample window functions.
  • When you omit the resolution ([1h:]), it is 1 minute.
  • APM calculates <expr> at the times that are multiples of the resolution in unix time. These times do not depend on the query start, so a dashboard refresh does not change the value at the same time.
  • The value at time t comes from the calculated values in (t − range, t]. With offset, the window is (t − offset − range, t − offset].
  • A selector in <expr> uses the last sample in the 5 minutes (τ − 5m, τ] before each calculation time τ (the same as Prometheus). When the resolution is shorter than the collection interval, the same sample value repeats. When there is no sample for more than 5 minutes, that time has no value. This is different from the range query value of a selector without a function (a step window aggregate).
  • A time without a value is not put in the window. It is not filled with 0.
  • You can put a subquery in a subquery (max_over_time(sum_over_time(x[30s:10s])[5m:10s])).
  • The result of absent_over_time(<subquery>) has the labels {} (the same as Prometheus). The other functions use the same label rules as the usual window functions.
  • One subquery can calculate up to 50,000,000 points (series × calculation times). Beyond that, the query gets a bad_data error (HTTP 400) subquery evaluates too many points (<points> > <limit>); use a coarser resolution. APM checks the number of times before the calculation and the number of points after it. You can change the limit with the collection server system property promql.subquery.max.points.
  • You cannot use a calculated metric (such as cpu_usage_percent) in a subquery.
  • APM does not approximate the values in a subquery. In these cases, the query gets a bad_data error (HTTP 400).
    • A selector in the subquery would read more than 22,000 small windows (promql.window.max.buckets) for one series, (query period + range + 5m) ÷ (greatest common divisor of 5m and the resolution). The error message is subquery needs too many buckets per series (<windows> > <limit>). APM checks this before the calculation starts. Use a larger resolution or a shorter range.
    • The greatest common divisor of 5 minutes and the resolution is shorter than 1 second (for example, [1m:700ms]).
    • A calculated metric is used in a subquery, or a window function in a subquery cannot be calculated exactly. The error message is subquery cannot be evaluated exactly: ….

Aggregation operators -- sum, avg, min, max, count, stddev, stdvar, topk(N, …), bottomk(N, …), quantile(0.95, …), count_values("label", …), group. Group with a by or without clause. A by result keeps only the listed labels. A without result removes the listed labels and __name__:

avg by (instance) (jvm_heap_heapUsed) # average per instance
sum without (user_key) (jvm_heap_heapUsed) # group by everything except user_key
topk(5, rate(gc_G1_Young_Generation_gcCount[5m])) # the top 5 instances by GC frequency
count(count by (ip_addr) (jvm_heap_heapUsed)) # number of hosts
count_values("heap_gib", round(jvm_heap_heapMax / 1024 / 1024 / 1024)) # number of instances for each maximum heap size (GiB)

The N of topk and bottomk and the value of quantile are scalar expressions (2 * 3, scalar(x)). In a range query, a value that differs per step is used at each step. N is rounded down. If N is less than 1 or NaN, that step has no result. A vector gives a bad_data error. Wrap it in scalar(). count_values writes the label value in the same format as Prometheus (1, 0.5, 1000000, NaN, +Inf).

Time shift -- offset and @:

jvm_heap_heapUsed offset 1h # the value one hour ago
rate(gc_G1_Young_Generation_gcCount[5m] offset 1d) # the GC frequency at the same time one day ago
jvm_heap_heapUsed @ 1700000000 # the value at a fixed time (unix seconds)
jvm_heap_heapUsed @ end() # the value at the end of the query range

Write offset directly after a selector, a range selector, or a subquery. After a function result or an expression in parentheses, offset gives a bad_data error. A negative offset (offset -5m) moves forward in time. @ reads the value at that one time. In a range query, all steps have the same value. @ start() and @ end() are the start and end of a range query. In an instant query, both are the evaluation time. You can use offset and @ together, in any order.

Label functions -- label_replace and label_join:

# the part of instance_id before "-" as the app label
label_replace(jvm_heap_heapUsed, "app", "$1", "instance_id", "(.*)-.*")
# ip_addr and instance_id joined with ":" as the host label
label_join(jvm_heap_heapUsed, "host", ":", "ip_addr", "instance_id")

The regular expression of label_replace uses the same RE2 syntax as label matchers and must match the full source label value. A series that does not match stays the same. In the replacement, use $1, ${1}, $name, ${name}, and $$. Define a named group as (?P<name>…). If the result is an empty string, the label is removed. If the label_replace result has two or more series with the same labels, you get a bad_data error.

Maths, arithmetic, and comparison:

abs(metric) ceil(metric) floor(metric) round(metric) round(metric, 0.5)
clamp(metric, 0, 100) clamp_min(metric, 0) clamp_max(metric, 100)
sqrt(metric) exp(metric) ln(metric) log2(metric) log10(metric) sgn(metric)
sin(metric) cos(metric) tan(metric) asin(metric) acos(metric) atan(metric)
sinh(metric) cosh(metric) tanh(metric) asinh(metric) acosh(metric) atanh(metric)
deg(metric) rad(metric) pi()

jvm_heap_heapUsed / jvm_heap_heapMax
jvm_heap_heapUsed / on(instance) cpu_usage_cpuUsage
rate(metric[5m]) * 60
jvm_heap_heapUsed > 1024 * 1024 * 1024 # returns only the series over 1 GB
jvm_heap_heapUsed > bool 1024 * 1024 * 1024 # 1 (over) or 0 for each series

The result of an arithmetic operation, a maths function, rate, irate, increase, *_over_time, or a raw sample window function does not have the __name__ label (last_over_time keeps it). In a range query, topk and bottomk rank the series again at every step. A series has only the points of the steps where it was picked, so a line in a graph can have gaps (the same as Prometheus). A comparison filter (without bool) keeps only the series that meet the condition, with their labels. Division by zero gives the value +Inf, -Inf, or NaN. A maths function given a value outside its domain keeps the series and gives NaN or -Inf (sqrt(-1) = NaN, ln(0) = -Inf). pi() is the scalar π.

Sorting and checking for series -- sort, sort_desc, absent:

sort_desc(jvm_heap_heapUsed) # largest value first
absent(jvm_heap_heapUsed{instance_id="api-53-x"}) # {instance_id="api-53-x"} 1 when the series does not exist
absent_over_time(jvm_heap_heapUsed{instance_id="api-53-x"}[10m]) # 1 when there is no value for 10 minutes

sort and sort_desc order the result of an instant query by value. NaN comes last, and equal values are ordered by labels. A range query keeps its order (the same as Prometheus). absent returns 1 at each time where no series of the argument has a value. When a value exists, the result is empty. The result labels come from the = matchers of the argument selector (__name__ excluded, app_name included). A label with two or more matchers is left out, and an argument that is not a selector (absent(sum(…))) gives no labels.

Time and scalar conversion -- time, scalar, vector, date functions:

time() # evaluation time (unix seconds); the step time in a range query
time() - timestamp(jvm_heap_heapUsed) # seconds since the last collected sample
scalar(sum(jvm_heap_heapUsed)) # a vector with one series as a scalar
jvm_heap_heapUsed > scalar(avg(jvm_heap_heapUsed)) # series above the overall average at each step
topk(scalar(count(jvm_heap_heapUsed) / 7), jvm_heap_heapUsed) # top 1/7 of the instances
vector(time()) # a scalar as a series without labels
hour() # hour of the evaluation time (0-23, KST by default)
year(vector(1700000000)) # year of unix time 1700000000 = 2023
  • time() is the evaluation time in unix seconds. In a range query it is the time of each step.
  • scalar(v) is, at each step, the value of the only series of v with a value there. With no series or two or more, it is NaN.
  • vector(s) is one series without labels whose value at each step is the value of s.
  • The date functions use KST (Asia/Seoul) by default: minute (0-59), hour (0-23), day_of_week (0 = Sunday to 6), day_of_month (1-31), day_of_year (1-366), days_in_month (28-31), month (1-12), year. hour() is 9 at 9 a.m. KST.
  • Set the time zone with the JVM system property promql.timezone on the collection server. The value is a time zone name such as Asia/Seoul or UTC (for example -Dpromql.timezone=UTC). An invalid value is logged as a warning and KST is used.
  • Prometheus uses UTC, so with the default setting the result differs from Prometheus by 9 hours. Set promql.timezone=UTC when you need the Prometheus value. observ (Cache PromQL) also defaults to KST.
  • Without an argument, a date function gives one series without labels computed from the evaluation time (the step time in a range query). With a vector argument, each value is read as unix seconds, and the result drops __name__. Fractions of a second are dropped, and NaN and ±Inf stay as they are.
  • A scalar can have a different value at each step. Every place that takes a scalar (arithmetic and comparison operators, round, clamp, clamp_min, clamp_max, the seconds of predict_linear, the φ of quantile_over_time, the parameter of topk, bottomk, quantile) uses the value of each step. A step where the minimum of clamp is greater than the maximum has no result.
  • When the query result is a scalar, an instant query returns result type scalar, and a range query returns one series without labels.
  • When v is a selector (also in parentheses), timestamp(v) is the time of the stored raw sample. At each time it returns the time of the last raw sample in the last 5 minutes ((t − offset − 5m, t − offset]) in unix seconds. Milliseconds are the decimal part. When there is no sample in 5 minutes, that time has no point (the same as Prometheus).
  • timestamp(v) reads raw samples, so the raw sample limit applies (see "Raw sample window functions" above). A long range query over the limit gets a bad_data error. With @, every time uses the same window. Inside a subquery the same rule applies at each inner calculation time.
  • When v is not a selector (timestamp(sum(x)), timestamp(rate(x[5m]))), it is the time of each point, which is the evaluation time (the step time in a range query). The result drops __name__. A range vector (timestamp(x[5m])) gets a bad_data error.

Logical operators -- and, or, unless:

jvm_heap_heapUsed and on(instance) jvm_heap_heapMax > 1024 * 1024 * 1024 # heap used of the instances whose maximum heap is over 1 GB
jvm_heap_heapUsed unless jvm_heap_heapUsed{instance_id="khan11"} # all except khan11
rate(gc_G1_Young_Generation_gcCount[5m]) > 1 or rate(gc_G1_Young_Generation_gcCount[5m]) < 0.1 # series that meet either condition
  • a and b keeps only the points of a for which b has a point with the same matching key at the same time. The values and labels are those of a.
  • a unless b keeps only the points of a for which b has no point with the same matching key at the same time.
  • a or b gives all points of a, plus the points of b for which a has no point with the same matching key at the same time. Each keeps its own labels.

The matching key is the same as for arithmetic operators: all labels except __name__ by default, changed with on(...) or ignoring(...). The result keeps __name__. A range query decides at each time separately. A scalar operand, or use with bool, group_left, or group_right, gives a bad_data error.

Many-to-one matching -- group_left, group_right:

jvm_heap_heapUsed / on(ip_addr) group_left cpu_usage_cpuUsage # several instances on one host
jvm_heap_heapUsed / on(instance) group_left(app_name) jvm_heap_heapMax # copies the app_name label of the right side to the result

Use group_left when the left side has several series with the same matching key and the right side has one. group_right is the reverse. It comes after on(...) or ignoring(...), and works only with arithmetic and comparison operators.

  • The result labels are the labels of the side with several series (the left side for group_left). Unlike one-to-one matching, they are not reduced to the on(...) labels. Arithmetic operators and bool comparisons remove __name__.
  • Each label listed in the parentheses is copied from the side with one series. When that side does not have the label, it is removed from the result.
  • When the side that must have one series has two or more series with the same matching key at the same time, you get a bad_data error (many-to-many matching not allowed). Two or more result series with the same labels also give an error.
  • The value of a comparison filter (without bool) is the value of the left operand (the same as Prometheus).
  • To put the right operand in parentheses, write an empty label list first, as in group_left() (a + b). Prometheus reads parentheses right after group_left as the label list.

Operator precedence -- an operator higher in the table binds first (the same as Prometheus).

OrderOperatorsGrouping
1^Right to left (2 ^ 3 ^ 2 = 2 ^ (3 ^ 2) = 512)
2Unary +, -(-2 ^ 2 = -(2 ^ 2) = -4)
3*, /, %, atan2Left to right
4+, -Left to right
5==, !=, <=, <, >=, >Left to right
6and, unlessLeft to right
7orLeft to right

Versions before 2026-10-05 calculated -2 ^ 2 as 4 and 2 ^ 3 ^ 2 as 64. Dashboards that use such expressions show different values.

a atan2 b is the arc tangent of a over b for each point (radians, −π to π). As with the arithmetic operators, you can use on(...), ignoring(...), group_left, and group_right, and the result drops __name__. With bool it is a syntax error (bad_data).


Metric Explorer​

A UI for running PromQL interactively in a browser -- reach it at the address in How to open it above.

AreaFunction
Metric selection dropdownAutomatic discovery of registered metrics (by __name__)
PromQL query inputType it directly -- choosing from the dropdown fills it in automatically
Label filter panelDropdowns per ip_addr, agent_type, instance_id, and namespace
Aggregate function selectionsum, avg, topk, quantile, and others
Time rangeQuick buttons from the last 5m to 7d, plus a custom range
StepAuto / 15s / 1m / 5m / 10m / 30m / 1h
Result visualizationLine, area, and bar charts plus a table of raw values
Copy API URLCopies the current query as an /api/v1/query_range URL -- for external integration such as Grafana

Making your first chart in 30 seconds:

  1. Select jvm_heap_heapUsed in the metric dropdown
  2. Select agent_type=WAS in the label filter
  3. Click the 1h time range and leave Step at Auto
  4. Click Run → a graph of one hour of heap usage across the WAS instances appears

Note Auto is usually enough for Step. A small step with a long range loads the server heavily.

Grafana integration -- the same API can be connected as a Prometheus data source in Grafana:

URL: http://{apm-server}/monitoring
Access: Server (default)

Grafana calls /api/v1/query, /api/v1/query_range, /api/v1/labels, and the rest automatically -- reuse existing dashboard and alerting assets as they are.


HTTP API Summary​

MethodPathPurpose
GET / POST/api/v1/queryInstant (single point in time) query
GET / POST/api/v1/query_rangeRange (time range) query
GET/api/v1/seriesSeries metadata
GET/api/v1/labelsThe list of label names
GET/api/v1/label/{name}/valuesThe list of values for a particular label
GET/api/v1/metadataMetric metadata
GET/api/v1/metric-explorerThe Metric Explorer UI

Parameters -- query (required), time for instant, and start, end, and step for range (calculated automatically when omitted). /series, /labels, and /label/{name}/values accept match[], start, end, and limit. With match[], only the label names and values of the series that match the selector are returned (the Grafana variable label_values(metric, label) uses this call). limit is the number of entries kept from the start of the sorted result (0 or omitted means no limit; a negative value is bad_data). Exploratory calls:

# The full list of available metrics (~250)
curl "http://{apm-server}/monitoring/api/v1/label/__name__/values"

# The list of label names / the values of a particular label
curl "http://{apm-server}/monitoring/api/v1/labels"
curl "http://{apm-server}/monitoring/api/v1/label/agent_type/values"

# The labels of one metric / 10 instance values of that metric
curl -G "http://{apm-server}/monitoring/api/v1/labels" --data-urlencode "match[]=jvm_heap_heapUsed"
curl -G "http://{apm-server}/monitoring/api/v1/label/instance/values" \
--data-urlencode "match[]=jvm_heap_heapUsed" --data-urlencode "limit=10"

When match[] has several selectors and the storage query of one of them fails, the whole request is an error, with no partial result (the same as Prometheus).

HTTP error status -- the errorType and HTTP status of an error response are the same as in Prometheus, except unavailable. All APIs use the same rule.

errorTypeHTTPCause
bad_data400Syntax error, unsupported function, query size limit exceeded, incorrect parameter
execution422Other execution errors (internal error of the collection server)
internal500The storage (InfluxDB) answered with an error. The message starts with InfluxDB query failed: …
timeout503The storage did not answer within the time limit, or the turn of a heavy query did not come within 10 seconds
unavailable503The collection server could not connect to the storage. Prometheus uses 500 for this errorType. APM uses 503 because you can try again later

For a 503, try again a little later.


A Collection of Frequently Used Queries​

The metric names are examples following the general pattern -- check the exact names for your environment in the Metric Explorer.

JVM / WAS:

jvm_heap_heapUsed{agent_type="WAS"} # current heap usage across all WAS
avg_over_time(jvm_heap_heapUsed[5m]) # 5-minute average heap usage
topk(5, jvm_heap_heapUsed) # the top 5 instances by heap usage
(jvm_heap_heapUsed / jvm_heap_heapMax) > 0.8 # heap utilization over 80%
sum_over_time(gc_G1_Young_Generation_gcTime[1h]) # total GC time (1 hour)
rate(gc_G1_Young_Generation_gcCount[5m]) # GC count per second
datasource_pool_active{agent_type="WAS"} # DataSource connections in use

System:

topk(10, cpu_usage_cpuUsage{agent_type="SYS"}) # the top 10 hosts by CPU
cpu_usage_systemCpuUsage # system-wide CPU
avg_over_time(cpu_usage_cpuUsage[1h]) # 1-hour average CPU per instance

Aggregating per host:

avg by (ip_addr) (jvm_heap_heapUsed) # average heap per host
count by (agent_type) (jvm_heap_heapUsed) # instance count per agent_type
jvm_heap_heapUsed{ip_addr="192.168.80.190"} # heap usage for a particular IP

When It Does Not Work​

SymptomWhat to check
You do not know the metric nameCheck the actual name in the Metric Explorer dropdown or at /api/v1/label/__name__/values
The result is emptyA typo in a label value -- check the possible values at /api/v1/label/{name}/values, and check the case of agent_type
Metric Explorer sign-in failsAn API access key is required -- create and copy the key under Settings ▸ Users
A rate() value is not what you expectThe calculation depends on how the field is stored -- see the storage table in "How It Differs from Prometheus" above. For a field already stored as a rate, a query without a function (the mean of the last 5 minutes) and rate(x[N]) (the mean of the last N) give different values
The response warnings has … is stored as a rate … (rate_alias)The field is already stored as a rate, so rate was calculated as avg_over_time and irate as last_over_time -- use avg_over_time or last_over_time directly to get no warning. The tps fields are calculated from count
The response warnings has the range is shorter than … (rate_short_window)The range is shorter than the 2-second stored interval -- set a range of 2 seconds or more
The sum or average does not matchA query without a function gives the mean or sum of the last 5 minutes (instant query) or of one step from each time (range query) -- write the window you want with sum_over_time or avg_over_time
The response warnings has window function … approximated with step bucketsA window function was calculated from one step window -- set a step that divides the range or query a shorter period (see "The calculation window" above)
query selects too many series errorOne selector or window function reads more series than the limit (2,000,000 ÷ rows per series) -- use a larger step, shorten the query period, or narrow the series with labels (see "Querying a Specified Period" above)
query reads too many rows errorThe rows one query reads in total are more than the limit (8,000,000 by default) -- use a larger step, shorten the query period, or narrow the series with labels
too many concurrent heavy queries; retry later error (HTTP 503)4 queries with more than 5,000 rows per series are running, and the turn did not come within 10 seconds -- run the query again a little later, or specify a larger step
internal error (HTTP 500) InfluxDB query failed: …The storage answered with an error -- check the InfluxDB error text in the message together with the InfluxQL in the collection server log
timeout or unavailable error (HTTP 503)The storage did not answer within the time limit, or it could not be reached -- check the InfluxDB status and run the query again a little later. If timeouts continue, shorten the range or specify a larger step
query reads too many samples errorA calculation that reads raw samples (a raw sample window function, the rate family calculated from raw samples, timestamp()) reads more samples than the limit (1,000,000 by default) -- shorten the range or the query period, or narrow the series with labels
subquery evaluates too many points errorA subquery calculates more points than the limit (50,000,000 by default) -- use a larger resolution, or shorten the query period or the range
subquery needs too many buckets per series errorA selector in the subquery reads more small windows than the limit per series (22,000 by default) -- use a larger resolution, or shorten the range or the query period
cannot be evaluated exactly error (subquery … cannot be evaluated exactly)APM cannot calculate the values in the subquery exactly at the resolution -- use a larger resolution or a shorter range. A calculated metric cannot be used in a subquery (see "Subqueries" above)
An operation between two metrics is emptyThe labels are different, so the series do not match -- specify the labels to compare, as in on(instance)
bad_data error (HTTP 400)Unsupported syntax or function -- check the function name in the error message and the "Not supported" table above
histogram_quantile does not workNot supported -- APM does not store le buckets. Query response-time fields such as apdex_avgRT and apdex_maxRT directly
A long-range query is slowSpecify a large step (30m to 1h) -- a small step over a long range is heavy

  • OPENMARU APM User Guide R2. Chart Metric Reference -- what each metric means, the signals of trouble, and thresholds
  • OPENMARU APM User Guide E3. What the Metrics Mean -- the concepts of APDEX, response time, and throughput
  • OPENMARU APM User Guide H3. My Dashboard -- gathering charts inside the console (no PromQL needed)
  • OPENMARU APM User Guide H18. Managing Users, Groups, and Permissions -- where the API access key is issued