<< All versions

Skill v0.1.0

currentAutomated scan100/100
radra23/otel-as-code-plugin/terraform-patterns
──Details
PublishedSeptember 28, 2026 at 12:23 AM
Content Hashsha256:274016a7e0861787...
Git SHA
──Files
Files (1 file, 31.0 KB)
SKILL.md31.0 KBactive
SKILL.md · 462 lines · 31.0 KB

name: terraform-patterns description: Per-backend Terraform provider patterns, resource names, auth variables, and gotchas for grafana / datadog / newrelic / dash0. Use when generating observability-backend Terraform. version: 0.1.0


Terraform Patterns for Observability Backends

What this file is, and what outranks it

This is a cache of provider schemas, hand-written and pinned to a provider major. It carries what a schema dump cannot — which resources are worth emitting, which queries are right for OTel data, and the per-vendor gotchas below. That is its value.

It is also, unavoidably, stale the moment a provider ships a release. So when a schema lookup is available (a Terraform registry/MCP tool, or a local terraform providers schema -json — see "Optional: verify against the live provider schema" in agents/terraform-gen.md), the live schema wins on facts: whether a resource exists, its required arguments, whether an argument is deprecated. This file keeps governing judgement.

Two consequences worth stating plainly:

  • A disagreement between this file and the live schema is a bug here, and should be reported

so it gets fixed — not silently worked around in generated output.

  • No lookup being available is normal. The gotchas below are still correct enough to generate

from, and the golden snapshots are validated offline in CI precisely so this file stays usable on its own.

Module Shape (all backends)

Every generated module has exactly three files:

  • main.tf — all resources + provider block
  • variables.tf — all input variables
  • outputs.tf — key output values (dashboard URL, monitor IDs, SLO IDs)

Header comment required in every main.tf:

hcl
# Generated by otel-as-code v0.1.0 on YYYY-MM-DD.
# Re-run /otel-backend <vendor> to regenerate.
# Drift detection: v2 roadmap.

Run terraform fmt and terraform validate immediately after writing the files. Emit the exact commands the user should run next (init, plan, apply) as a post-generation note.

Service name → resource identifiers (sanitize, but ONLY in identifier positions)

service.name is frequently an npm-scoped package name like @myorg/web (from package.json#name, confidence 0.97) — @ and / are invalid in most resource identifiers. terraform validate does NOT catch this (the invalid string is inside an embedded YAML / a uid the provider accepts as an opaque string), so it only surfaces at apply against a live account. Derive a sanitized slug and use it in identifier positions:

hcl
locals {
service_slug = trim(replace(lower(var.service_name), "/[^a-z0-9]+/", "_"), "_")
}
  • Sanitize (use `local.service_slug`) wherever the name becomes an identifier: a Grafana

dashboard uid / rule-group name, a Prometheus/Dash0 alert: name, a Kubernetes metadata.name, any slug. Pick the separator the target format allows: _ for Prometheus alert names and Grafana uids (both permit _); - for a DNS-1123 metadata.name (which forbids _) — the two constraints have no common separator, so choose per field, don't assume one slug fits all.

  • Do NOT sanitize the value used in a query/filter ({service_name="..."},

{job="..."}, WHERE service.name = '...') — it must equal the emitted service.name, so it keeps the raw var.service_name; a slug there silently matches nothing.

  • Do NOT sanitize a display title (name = "High error rate — ${var.service_name}") —

free text, the real name reads better.

Workloads without an OTel SDK (scraped metrics)

Some in-scope services emit no OTel SDK telemetry, such as an upstream binary like Keycloak (the scanner marks these generatorSupported: false, inScope: true). Their metrics come from a Prometheus endpoint that the Collector scrapes, so the OTel semconv names used in the backend sections below (http_server_request_duration_seconds, http_response_status_code) don't exist for them. Use the names in this section instead. Don't guess other names, and don't leave a panel's query empty: put "Verify against your scrape config: <metric>" in the description of every panel and alert built from this table.

Two rules differ from the OTel-SDK queries elsewhere in this file:

  • `job` is the scrape job, not the service name. The Collector's prometheus receiver sets

service.name from the scrape config's job_name, and that becomes the job label. Filter by the job_name in the user's scrape config (read it from the Collector config in the repo, or ask). It is often not the scanner's service name. Three variations:

  • If the Collector sets service.namespace, job is <namespace>/<job_name>.
  • If the Collector re-exposes metrics through its prometheus exporter and a Prometheus server

scrapes that, the server's own job wins unless its scrape config sets honor_labels: true.

  • On Dash0, filter {service_name="<scrape job>"} instead (see the Dash0 section below).
  • Names follow the exporter, not OTel semconv. Micrometer (Quarkus, Spring Boot, Keycloak)

names its HTTP server timer http_server_requests_seconds_* with labels method, uri, status and outcome.

These are the names on the scrape endpoint. PromQL backends (Grafana, Dash0) normally see them unchanged after a Collector prometheus receiver, unless the receiver sets trim_metric_suffixes: true, which drops unit and type suffixes (http_server_requests_seconds becomes http_server_requests, node_cpu_seconds_total becomes node_cpu). Check the Collector config. Datadog and New Relic apply their own ingestion naming, which this file doesn't catalog: look the metric up in the account before writing a query for those backends.

SourceMetricWhat it isQuery notes
Micrometer HTTP server (Quarkus, Spring Boot, Keycloak)http_server_requests_seconds_count, _sum, _maxrequest count, total seconds, max seconds; labels method, uri, status, outcomeRequest rate sum(rate(http_server_requests_seconds_count{job="<scrape job>"}[5m])); errors add status=~"5..". _bucket exists only when histograms are enabled (Keycloak: http-metrics-histograms-enabled=true); without it, mean latency is sum(rate(http_server_requests_seconds_sum{job="<scrape job>"}[5m])) / sum(rate(http_server_requests_seconds_count{job="<scrape job>"}[5m])) (never a raw _sum / _count, which averages since process start and barely moves), and max(http_server_requests_seconds_max{job="<scrape job>"}) is a decaying worst case, not a quantile. Never histogram_quantile without _bucket
Micrometer JVMjvm_memory_used_bytes, jvm_memory_committed_bytes, jvm_gc_pause_seconds_count / _sum / _maxheap and non-heap memory; GC pausesGC pause share: sum by (instance) (rate(jvm_gc_pause_seconds_sum{job="<scrape job>"}[5m])) (series are split by action, cause and gc, so aggregate). Concurrent collectors (ZGC, Shenandoah) do most work outside pauses; Keycloak also exposes jvm_gc_overhead directly
Keycloak user eventskeycloak_user_events_totalcounter per user event; labels realm, event (e.g. login, logout), error (empty string on success); client_id and idp are off by defaultOff by default: needs --metrics-enabled=true --event-metrics-user-enabled=true. Login rate sum(rate(keycloak_user_events_total{job="<scrape job>",event="login",error=""}[5m])); failures use error!=""
blackbox_exporterprobe_success, probe_duration_seconds, probe_http_status_code, probe_ssl_earliest_cert_expiryprobe result (1/0), duration, status code, earliest certificate expiry as a Unix timestamp in secondsDays to cert expiry (probe_ssl_earliest_cert_expiry{job="<scrape job>"} - time()) / 86400; alert when it drops below the renewal lead time, e.g. < 14
node_exporternode_cpu_seconds_total (labels cpu, mode), node_memory_MemAvailable_bytes, node_filesystem_avail_bytes, node_filesystem_size_byteshost CPU, memory, diskCPU busy per host 1 - avg by (instance) (rate(node_cpu_seconds_total{job="<scrape job>",mode="idle"}[5m])) (a bare avg() hides one saturated host among idle ones); disk free share `node_filesystem_avail_bytes{job="<scrape job>",fstype!~"tmpfsoverlaysquashfs"} / node_filesystem_size_bytes{job="<scrape job>",fstype!~"tmpfsoverlaysquashfs"}` (pseudo filesystems read near 0% or 100%)

Verified against upstream source, not memory: Quarkus telemetry-micrometer.adoc; Keycloak docs/guides/observability/metrics-for-troubleshooting-{http,jvm,keycloak}.adoc and event-metrics.adoc; blackbox_exporter prober/; node_exporter collector/; the Collector prometheusreceiver README for the job_name → service.name mapping. If an exporter is not in this table, find its metric names in its own docs and say which source you used.


Grafana Cloud (grafana/grafana ~> 4.0)

Authentication variables

hcl
variable "grafana_url" {
description = "Grafana Cloud URL (e.g. https://yourorg.grafana.net)"
type = string
}
variable "grafana_service_account_token" {
description = "Service account token with Editor role"
type = string
sensitive = true
}

Required resources

  1. grafana_folder — create a folder to scope all generated resources
  2. grafana_dashboard — uses config_json with a JSON-encoded dashboard model
  3. grafana_rule_group — unified alerting (NOT the deprecated grafana_alert_notification)
  4. grafana_slo — Grafana Cloud SLOs (requires grafana_slo resource, available on Cloud plans)

Key gotchas

  • grafana_folder.uid is auto-generated on create; reference it via grafana_folder.<resource_label>.uid (e.g. grafana_folder.otel_folder.uid)
  • Dashboard config_json must be valid Grafana JSON; use jsonencode() to construct it safely
  • grafana_rule_group requires a folder_uid (from the folder resource) and interval_seconds
  • SLO query block uses PromQL expressions; prefer rate() over irate() for stability
  • Pin ~> 4.0 (current major); the 2.x → 3.x → 4.x jumps each changed resource schemas — don't copy older-major patterns
  • Filter by `job`, NOT `service_name`. On a default OTLP → Prometheus pipeline, service.name is mapped to the `job` label (and service.instance.id → instance); resource attributes are NOT added as metric labels — they live only on the separate target_info series (otlp.promote_resource_attributes defaults to []). So {service_name="..."} returns NO DATA — a silently-empty panel/alert. Use {job="<name>"}; when service.namespace is set, job becomes <namespace>/<name>, so filter {job="<namespace>/<name>"}. (Only emit {service_name="..."} if the user's Prometheus receiver sets otlp.promote_resource_attributes: [service.name].)

OTel-specific dashboard panels to generate

  • Request rate: rate(http_server_request_duration_seconds_count{job="..."}[5m])
  • Error rate: rate(http_server_request_duration_seconds_count{job="...",http_response_status_code=~"5.."}[5m])
  • P99 latency: histogram_quantile(0.99, rate(http_server_request_duration_seconds_bucket{job="..."}[5m]))

Business-attribute panels (from confirmed businessAttrs)

Grafana panels query Prometheus, so a business attribute is only visible as a metric (or a metric label). Sanitize the attribute name to a Prometheus metric name (dots/dashes → underscores) as <M>, and scope every query to the service by job — exactly as the generic panels do (service.name → job; <namespace>/<name> when a namespace is set). Emit the caveat as the panel description so an empty panel is self-explanatory.

  • kind: counter → sum(rate(<M>_total{job="<name>"}[5m])) (OTLP→Prometheus appends _total to

counters). Caveat: requires the app to emit a counter named <M>.

  • kind: gauge → avg(<M>{job="<name>"}) (or avg_over_time(<M>{job="<name>"}[5m])) — plot the

value directly. Do NOT wrap a gauge in rate()/sum(rate()): rate() is defined only for counters, and summing a ratio like conversion_rate is meaningless.

  • kind: dimension → sum by (<label>) (rate(http_server_request_duration_seconds_count{job="<name>"}[5m])).

Caveat: the breakdown needs <label> to be a metric label. If the attribute is recorded as a per-request metric data-point attribute, it already is one; if it is only a resource attribute, it is NOT a label on a default OTLP→Prometheus pipeline (see the job-label gotcha above) unless the producer promotes it (otlp.promote_resource_attributes).

  • kind: histogram → three query targets in one panel, one per quantile (OTLP→Prometheus emits a

native Prometheus histogram: <M>_bucket/<M>_sum/<M>_count): histogram_quantile(0.50, sum(rate(<M>_bucket{job="<name>"}[5m])) by (le)), histogram_quantile(0.95, ...), histogram_quantile(0.99, ...). Label the y-axis with the entry's unit when present. Caveat: requires the app to emit a histogram named <M>; never substitute avg(<M>_sum{job="<name>"} / <M>_count{job="<name>"}) for this — an average, not a quantile, defeats the reason a histogram was captured.

Notification routing (contact points)

Emit contact points only when the user asks for alert routing. Without one, the generated grafana_rule_group alerts go through the stack's existing notification policy.

  • Prefer a native integration. grafana_contact_point has blocks for many receivers (Slack,

PagerDuty, email, Opsgenie, and more). Use the matching block when one exists.

  • `webhook` sends Grafana's own JSON (receiver, status, alerts[], commonLabels,

title, message, ...). For a receiver that expects a different shape, in order of preference:

  1. The receiver reads Grafana's payload. ntfy v2.14.0 and later has a built-in grafana

template: set url to https://<ntfy-host>/<topic>?template=grafana and ntfy formats the title and message itself. No adapter is needed.

  1. Priority and other request options: ntfy takes priority from the X-Priority header or

the priority/p query parameter. Prefer the query parameter (?template=grafana&priority=5): it needs no headers and so no ~> 4.6 pin. Either way it is fixed per contact point. For priority by severity, create one contact point per level and route each rule to its own (see below), or, on ntfy v2.17.0 and later with a server the user controls, use a custom server-side template whose priority: reads .commonLabels.severity. Use headers = { ... } only for a receiver that accepts no query-parameter form.

  1. Custom body: payload { template = "..." vars = { ... } } replaces the whole body; the

title and message fields are then ignored. Grafana's docs say Custom Payload is not yet generally available in Grafana Cloud, so use it only when the user confirms their Grafana supports it.

  1. Otherwise the receiver needs a translating adapter between it and Grafana. The module

can't provide one: say so in a comment on the contact point.

  • Pin `~> 4.6` when using `headers` or `payload`. Both first appear in grafana/grafana v4.6.0;

~> 4.0 would accept an older 4.x release that rejects them.

  • Route per rule, not with `grafana_notification_policy`. That resource manages the entire

notification policy tree and overwrites policies it didn't create, including the user's own routing. Set notification_settings { contact_point = grafana_contact_point.<label>.name } on each rule in the grafana_rule_group instead. This needs Grafana 10.4 or later; on 10.4.x the alertingSimplifiedRouting feature flag must be enabled (it is on by default from Grafana 11.0). State that in a comment.

  • Keep the URL out of the module. A contact point URL often embeds a secret (a private ntfy

topic, a tokened endpoint). Pass it from a sensitive = true variable.

Verified against the grafana/grafana provider docs at v4.46.0 (contact_point, rule_group, notification_policy), the provider docs at each 4.x tag (to find v4.6.0), Grafana's webhook notifier docs, Grafana's featuremgmt/registry.go at v10.4.0 and v11.0.0, and ntfy's publish.md and releases.md.


Datadog (DataDog/datadog ~> 4.0)

Authentication variables

hcl
variable "datadog_api_key" {
description = "Datadog API key"
type = string
sensitive = true
}
variable "datadog_app_key" {
description = "Datadog application key"
type = string
sensitive = true
}
variable "datadog_site" {
description = "Datadog site (e.g. datadoghq.com, datadoghq.eu)"
type = string
default = "datadoghq.com"
}

Required resources

  1. datadog_dashboard — use widget blocks (not JSON string)
  2. datadog_monitor — type "metric alert" for threshold-based; "query alert" for formula-based
  3. datadog_service_level_objective — type "metric" for metric-based SLOs

Key gotchas

  • Both api_key AND app_key are required; setting only one will silently fail
  • datadog_monitor.type must be one of: "metric alert", "service check", "event alert", "query alert", "composite", "log alert", "rum alert", "trace-analytics alert"
  • SLO timeframe must be exactly "7d", "30d", or "90d" — no other values accepted
  • Monitor tags must include the service tag: "service:<service_name>"
  • Dashboard layout_type is either "ordered" or "free"

OTel-specific monitor queries

Datadog APM trace metrics are named trace.<operation_name>.{hits,errors,duration}. For an OTel HTTP server span, Datadog's operation-name logic v2 (default on OTel Collector >= v0.126.0 / Datadog Agent >= v7.65) assigns the operation name http.server.request — NOT http.request. Use the trace.http.server.request stem:

  • Request rate: sum(last_5m):sum:trace.http.server.request.hits{service:<name>}.as_rate()
  • Error rate: sum(last_5m):sum:trace.http.server.request.errors{service:<name>}.as_rate() / sum:trace.http.server.request.hits{service:<name>}.as_rate() > 0.05
  • P99 latency: avg(last_5m):p99:trace.http.server.request{service:<name>} > 500000000

Two gotchas baked into those queries:

  • Percentiles need the bare distribution metric. trace.<op>.duration is a COUNT (total time) and does NOT support p50/p95/p99. Query the suffix-less distribution metric trace.http.server.request for percentile aggregations.
  • Durations are in nanoseconds. 500000000 = 500 ms. Do not write > 500 or > 0.5.
  • The operation name is pipeline-dependent. On pre-v2 collectors, or where a transform processor sets operation.name / legacy span_name_as_resource_name is configured, the stem differs. Emit a comment telling the user to confirm theirs in Datadog under APM > Metrics (Metrics Explorer, search trace.) before trusting the monitors — a monitor on a metric that never populates silently never fires.

Business-attribute panels (from confirmed businessAttrs)

  • kind: counter → sum:<name>{service:<name>}.as_rate(). Caveat: requires the counter reported

to Datadog; .as_rate() is documented for StatsD/DogStatsD rate/count metrics — for an OTLP-ingested counter, have the user confirm it populates, or drop .as_rate() for a plain sum:<name>{service:<name>}.

  • kind: gauge → avg:<name>{service:<name>} — the value as-is, no .as_rate().
  • kind: dimension → group the request metric by the tag:

sum:trace.http.server.request.hits{service:<name>} by {<tag>}.as_rate(). Caveat: the tag must be present on the spans/metrics (an OTel span attribute promoted to a Datadog tag).

  • kind: histogram → three query targets in one panel, one per percentile — Datadog distribution

metrics support a percentile aggregation prefix the same way the APM trace-metric queries above do (p99:trace.http.server.request{...}): p50:<name>{service:<name>}, p95:<name>{service:<name>}, p99:<name>{service:<name>}. Label the axis with the entry's unit when present. Caveat: requires the metric to actually be ingested as a Datadog distribution (not a gauge/count) — have the user confirm the metric type in Metrics Explorer if the panel is empty; an OTLP histogram maps to a Datadog distribution by default, but a custom exporter/pipeline could remap it.


New Relic (newrelic/newrelic ~> 3.0)

Authentication variables

hcl
variable "newrelic_account_id" {
description = "New Relic account ID"
type = number
}
variable "newrelic_api_key" {
description = "New Relic User API key (NRAK-...)"
type = string
sensitive = true
}
variable "newrelic_region" {
description = "New Relic region: US or EU"
type = string
default = "US"
}

Required resources

  1. newrelic_one_dashboard — pages with widget blocks; use widget_line and widget_table
  2. newrelic_nrql_alert_condition — NRQL-based alerts; type = "static" for threshold
  3. newrelic_service_level — SLOs; events block uses valid_events / good_events (and optionally bad_events) NRQL query blocks — NOT valid / good

Key gotchas

  • account_id has no provider-level default and must be set per resource — but WHERE differs by

resource. It is a top-level argument on newrelic_one_dashboard, newrelic_alert_policy, and newrelic_nrql_alert_condition. On newrelic_service_level there is no top-level account_id — it goes inside the events block (a top-level account_id there fails terraform validate with "An argument named account_id is not expected here"). Verified against newrelic/newrelic v3.x via terraform providers schema -json; re-check if you bump the pin.

  • newrelic_nrql_alert_condition requires an alert_policy_id; always create newrelic_alert_policy first
  • Service level good query denominator must return a rate between 0 and 1 — divide by total count
  • NRQL uses FROM Span for OTel trace data; attribute names follow OTel semconv directly
  • Dashboard widgets require account_id inside the nrql_query block, not just on the resource
  • Alert-condition NRQL must NOT contain `SINCE` / `UNTIL` / `TIMESERIES` / `COMPARE WITH` — New Relic rejects them in newrelic_nrql_alert_condition (the condition's own aggregation window drives timing). Those clauses are dashboard-only.

OTel-specific NRQL queries

Dashboard widget NRQL (a time window is expected — SINCE / TIMESERIES are fine here):

  • Request rate: SELECT rate(count(*), 1 MINUTE) FROM Span WHERE service.name = '<name>' SINCE 5 MINUTES AGO
  • Error rate: SELECT filter(count(*), WHERE otel.status_code = 'ERROR') / count(*) FROM Span WHERE service.name = '<name>' SINCE 5 MINUTES AGO
  • P99 latency: SELECT percentile(duration.ms, 99) FROM Span WHERE service.name = '<name>' SINCE 5 MINUTES AGO

Alert-condition NRQL (newrelic_nrql_alert_condition.nrql.query — NO SINCE/TIMESERIES):

  • Error rate: SELECT filter(count(*), WHERE otel.status_code = 'ERROR') / count(*) FROM Span WHERE service.name = '<name>'
  • P99 latency: SELECT percentile(duration.ms, 99) FROM Span WHERE service.name = '<name>'

Business-attribute widgets (New Relic can query spans/metrics directly — the strongest of the four for this). Backtick-quote dotted attribute names:

  • kind: counter → `SELECT sum(<name>) FROM Metric WHERE service.name = '<name>' TIMESERIES`

(a reported OTel counter; use `rate(sum(<name>), 1 minute)` for a per-minute rate).

  • kind: gauge → `SELECT average(<name>) FROM Metric WHERE service.name = '<name>' TIMESERIES`

(latest(...) for a level) — never sum() a ratio/level.

  • kind: dimension → `SELECT count(*) FROM Span WHERE service.name = '<name>' FACET <name> SINCE 5 MINUTES AGO` — a native breakdown of traffic by the business dimension (no metric-label caveat: the attribute is on the span).
  • kind: histogram → one widget, one NRQL query producing all three quantiles (NRQL's

percentile() accepts multiple values in one call, unlike the other three backends): `SELECT percentile(<name>, 50, 95, 99) FROM Metric WHERE service.name = '<name>' TIMESERIES. Label the axis with the entry's unit when present. Caveat: requires the histogram to be reported as a New Relic Metric; if it is only present as span/event data, query percentile(<name>, 50, 95, 99) FROM Span ` instead.


Dash0 (dash0hq/dash0)

Authentication variables

hcl
variable "dash0_auth_token" {
description = "Dash0 API auth token (must start with \"auth_\" or \"dash0_at_\"). Requires management/write API access, NOT an ingestion-only scope (this module creates dashboards/check rules) — a 403 on every resource means check the token's permission scope. Create one under Organization Settings > Auth Tokens."
type = string
sensitive = true
default = ""
}
variable "dash0_url" {
description = "Dash0 API endpoint (region-specific)"
type = string
default = "https://api.us-west-2.aws.dash0.com"
}
variable "dash0_dataset" {
description = "Dash0 dataset identifier (not display name). A data-partitioning concept, UNRELATED to the shared `environment` variable — never default this to \"production\" (not a dataset Dash0 provisions; produces a 403 that reads like a permissions failure but isn't). A fresh org's actual default dataset is named \"default\"."
type = string
default = "default"
}

Provider version

Use the latest available version from the Terraform registry. Check https://registry.terraform.io/providers/dash0hq/dash0/latest for the current version and add a version constraint to required_providers. Do not omit the version constraint.

Key gotchas

  • `check_rule_yaml` must be a full `PrometheusRule` document — never a flat alert body. It

needs apiVersion: monitoring.coreos.com/v1, kind: PrometheusRule, metadata.name, and spec.groups wrapping the rule; Dash0 currently supports exactly one group containing exactly one rule per check_rule_yaml. A flat alert: / expr: / for: / ... document with no groups wrapper fails live apply with error converting check rule YAML to Dash0 format: currently only one group is supported (confirmed via a live apply — the golden shipped with exactly this flat, broken shape until this was caught). Shape (mirror this, one rule per resource): ```yaml apiVersion: monitoring.coreos.com/v1 kind: PrometheusRule metadata: name: <dns-1123-safe-slug>-<alert-purpose> spec: groups:

  • name: Alerting

rules:

  • alert: <AlertName>

expr: <promql> for: 5m labels: { severity: critical } annotations: { summary: "..." } `` metadata.name follows **Kubernetes object-naming convention** (DNS-1123: lowercase alphanumeric + -, no underscores) — a separate constraint from the alert: field (a free-text Prometheus label value, where underscores are fine and were NOT the cause of the failure above). Sanitize service.name into a hyphen-based slug for metadata.name specifically (the same @scope/pkg problem as the "Service name → resource identifiers" rule above, but with - as the separator here since DNS-1123 forbids _`) — do not reuse an underscore-based slug for it.

  • `dash0_dataset` must never default to `"production"`. See the variable description above —

confirmed live: a nonexistent dataset produces a 403 that is easily misdiagnosed as a token-permission problem (#102 → #103). Default it to "default" (a fresh org's actual out-of-the-box dataset) and say so explicitly in the description, so a later regeneration does not reintroduce the mistake by pattern-matching onto the shared environment variable's "production" default — datasets and deployment environments are unrelated Dash0 concepts.

  • A `dash0.com/folder-path` annotation, if you add one, MUST start with a leading `/`. Dash0's

live API rejects an omitted leading slash with dash0 api error: folder path must start with '/' (status: 400) (confirmed via a live apply, #104) — "otel-as-code" fails, "/otel-as-code" succeeds. This is a fixed literal in the generated template, not account-specific, so it reproduces identically for every user if it regresses.

  • The auth token needs management/write scope, not ingestion-only. See the dash0_auth_token

description above — confirmed live: an insufficiently-scoped token produces a 403 on every resource (#102). Dash0's own docs distinguish ingestion-scoped tokens (send telemetry) from management-scoped ones (manage dashboards/check rules via the API this module uses); verify the exact current UI label for the broader scope against Dash0's own docs/account rather than asserting one here, since it was not independently confirmed.

Notes for implementors

Dash0 is an OTel-native backend; their Terraform provider models resources around OTel data directly. Before writing the terraform-gen subagent's Dash0 template, verify current resource names at https://registry.terraform.io/providers/dash0hq/dash0/latest/docs.

As of mid-2026, the provider supports dashboards and monitoring rules. Use the registry docs as the authoritative source for resource names and required fields — do not rely on this skill alone for Dash0-specific field names.

Key principle: Dash0 uses OTel attribute names natively in query expressions, so no translation layer is needed between OTel semconv and the monitoring query syntax.

Service filter: because Dash0 is OTel-native, its queries CAN filter by service directly — service.name in the Query Builder, or service_name (dots→underscores) in PromQL panels. This is unlike a vanilla OTLP→Prometheus pipeline, where service.name is only the job label and {service_name="..."} returns no data (see the Grafana gotcha above). So the Dash0 golden keeps {service_name="..."} intentionally — verify against a live Dash0 instance, as the exact PromQL label spelling depends on the panel/query surface.

Business-attribute panels (from confirmed businessAttrs)

Dash0 panels are PromQL, but its OTel-native pipeline makes business attributes far more likely to be queryable than a vanilla Prometheus setup (it filters by service_name directly). Sanitize the attribute name to a PromQL metric/label (dots→underscores) as <M>.

  • kind: counter → sum(rate(<M>_total{service_name="<name>"}[5m])). Caveat: requires the counter

to be emitted.

  • kind: gauge → avg(<M>{service_name="<name>"}) — the value as-is, not rated.
  • kind: dimension → sum by (<M>) (rate(http_server_request_duration_seconds_count{service_name="<name>"}[5m]))

— a breakdown by the business attribute, which Dash0's OTel-native ingestion keeps queryable as a label where a vanilla Prometheus pipeline would not. Verify the label spelling against a live instance.

  • kind: histogram → three query targets in one panel, one per quantile (same PromQL shape as

Grafana above — Dash0's PromQL surface exposes the same _bucket/_sum/_count histogram form): histogram_quantile(0.50, sum(rate(<M>_bucket{service_name="<name>"}[5m])) by (le)), histogram_quantile(0.95, ...), histogram_quantile(0.99, ...). Label the axis with the entry's unit when present. Caveat: requires the histogram to be emitted as <M>; verify against a live instance, as with the dimension case above.

All versions