Skip to main content
This feature is in and subject to change. To share feedback and/or issues, contact Support.
Active Session History (ASH) is a time-series sampling-based observability feature that helps you troubleshoot workload performance issues by capturing what work was actively executing on your cluster at specific points in time. Unlike traditional statement statistics that aggregate data over time, ASH provides point-in-time snapshots of active execution, making it easier to diagnose transient performance problems and understand resource usage patterns. ASH is accessible via CockroachDB SQL and is enabled by default. To enable or disable ASH, refer to Enable or disable Active Session History.

How ASH sampling works

ASH captures point-in-time snapshots of each ’s active work by sampling cluster activity at regular intervals (determined by the obs.ash.sample_interval cluster setting). At each sample point, ASH examines all goroutines that are actively executing or waiting, and it records what each one is doing. Each sample captures details like the workload that the goroutine belongs to (statement fingerprint, job ID, or system task) and the activity the goroutine is occupied with. These samples fill an in-memory circular ring buffer, the size of which is determined by the obs.ash.buffer_size cluster setting. When the buffer fills, the oldest samples are overwritten. Samples are also periodically flushed from memory to the system.active_session_history table (determined by the obs.ash.flush.enabled and obs.ash.flush.interval cluster settings). Persisted samples survive node restarts. A compaction job deletes persisted samples older than the retention period (determined by the obs.ash.compaction.retention_period cluster setting). The system.active_session_history table is excluded from . A user can query the in-memory ASH data for the specific node to which their SQL shell is connected (the gateway node) or across the whole cluster, as well as the persisted samples. Refer to ASH table reference. Because ASH is sampling-based rather than event-based, the sample count for a particular activity is proportional to how much time was spent on that activity. For example, if a query appears in 45 out of 60 sample points over one minute, it was actively consuming resources for approximately 45 seconds of that minute. This approach provides an accurate picture of resource usage patterns over time, but it means that short-lived operations completed between sampling points may not be captured. To troubleshoot very brief operations, you may need to reduce the obs.ash.sample_interval , or use statement diagnostics. However, use caution when reducing the sampling interval, as this will cause the buffer to fill up quickly. ASH does not provide exact resource accounting per query or job, nor does it provide the exact timing of individual executions. Its sampling-based approach instead provides a statistically reliable view of the system at a given time, which can help with troubleshooting. These point-in-time samples can be used to:
  • Root-cause slow queries: Understand exactly what a query was doing at specific points in time (e.g., , performing I/O, consuming CPU).
  • Identify bottlenecks: Determine which activities (such as lock waits, remote RPCs, admission control queues, or Raft proposals) are constraining workload performance.
  • Troubleshoot transient issues: Diagnose performance problems that don’t show up in aggregated statistics because they’re intermittent or short-lived.
  • Investigate past incidents: Reconstruct which workloads and wait events were active during an incident after it has ended, using persisted samples from the retention period.
  • Analyze resource usage patterns: Understand how different workloads (user queries, , system operations) consume cluster resources.
  • Compare performance across time: Analyze how workload behavior changes during different time periods (e.g., peak vs. off-peak hours).

Sample enrichment

Samples can be enriched with per-execution attributes: the user, plan_gist, canary_stats, txn_id, and session_id columns described in ASH table reference. Enrichment is disabled by default. To enable it, refer to Enable sample enrichment. When enrichment is enabled, per-execution attributes are cached on the gateway node. At each sample point, the sampler resolves a sample’s attributes from the local cache or from the sample’s gateway node. Samples that cannot be resolved immediately are retried at later sample points. If the attributes do not resolve within the enrichment window, the enrichment columns remain NULL. Enrichment behavior is configured with the obs.ash.enrichment.* cluster settings. In a cluster undergoing a to v26.3, enrichment columns remain NULL until the upgrade is finalized.

Use ASH alongside other monitoring tools

ASH complements CockroachDB’s existing observability tools. Each tool serves a different purpose: ASH can often be used with these other tools to help troubleshoot issues. For example, Prometheus metrics might alert you to a problem (such as a CPU spike at 2:15 PM). ASH shows which queries and jobs were actively running during that spike and which took a disproportionate amount of time to run. The Statements Page then provides aggregated performance data for those queries over a longer period of time, and statement diagnostics give detailed execution plans for deeper analysis.

Configuration

The following configure ASH behavior:

ASH table reference

ASH data is accessible through the following views in the system catalog:
  • : Includes in-memory samples from the gateway node.
  • : Includes in-memory samples from all nodes in the cluster. Querying the cluster-wide view may be more resource-intensive for large clusters.
  • : Includes persisted samples from all nodes in the cluster, retained for the retention period (determined by the obs.ash.compaction.retention_period cluster setting).
All of the views have the same columns:

query columns

For samples with a workload_type of STATEMENT, the query, query_summary, and database columns identify the sampled statement, resolved from its fingerprint. These columns are NULL for other workload types, and NULL until the statement’s fingerprint has been persisted to .

Enrichment columns

The user, plan_gist, canary_stats, txn_id, and session_id columns apply only to samples with a workload_type of STATEMENT, and are populated only when sample enrichment is enabled. They are NULL for other workload types, NULL when enrichment is disabled, and NULL for samples whose attributes were not resolved within the enrichment window.

workload columns

Each sample is attributed to a workload via the workload_type and workload_id columns. The encoding of workload_id depends on the workload_type:

work_event columns

The work_event column identifies the specific activity a goroutine was performing or waiting on when it was sampled, such as LockWait, DistSenderRemote, or KVEval. It is the most precise description of what a sample represents, and it is generally the column to group or filter by when analyzing ASH data. The work_event_type column is a coarse category associated with each work_event: CPU, IO, LOCK, NETWORK, ADMISSION, or OTHER. This category is a static, predefined mapping: every sample with a given work_event always has the same work_event_type, regardless of what the goroutine was actually doing at the time. As a result, work_event_type describes the resource that a work_event is generally concerned with, not the resource the sample was actually consuming. For example, a KVEval sample is categorized as IO, but the goroutine may have been using CPU.
Do not rely on work_event_type alone to determine whether a workload is CPU-bound, I/O-bound, or network-bound. Instead, look at the specific work_event values and the workloads they are attributed to, and correlate them with other signals such as and .
The following sections list the work_event values, grouped by their associated work_event_type.

CPU

work_events whose work_event_type is CPU are associated with active computation:

IO

work_events whose work_event_type is IO are associated with storage I/O:

LOCK

work_events whose work_event_type is LOCK are associated with lock and latch contention:

NETWORK

work_events whose work_event_type is NETWORK are associated with remote RPCs:

ADMISSION

work_events whose work_event_type is ADMISSION are associated with admission control queues:

OTHER

work_events whose work_event_type is OTHER are associated with miscellaneous wait points:

Debug zip integration

When the environment sampler triggers or , ASH writes aggregated report files (.txt and .json) alongside them. These reports are included in output. The naming pattern for these files is as follows:
  • TIMESTAMP: When the report was made (formatted as 2006-01-02T15_04_05.000)
  • TRIGGER: What event triggered the report (goroutine_dump or cpu_profile)
  • FORMAT: .txt (human-readable) or .json (structured)
For example: ash_report.2026-03-05T12_00_00.000.goroutine_dump.txt The lookback window for these reports is controlled by the obs.ash.log_interval cluster setting.

Common use cases and examples

ASH is accessed through the built-in CockroachDB SQL shell. Run to open the shell. CockroachDB Cloud deployments can also use the on the Console.

Enable or disable Active Session History

ASH sampling is enabled by default. To disable it on your cluster:
To re-enable it:
Enabling ASH begins collecting samples immediately. The in-memory buffer will fill up over time based on workload activity and the configured ASH cluster settings.

Enable sample enrichment

Sample enrichment is disabled by default. To populate the enrichment columns (user, plan_gist, canary_stats, txn_id, session_id) on new samples:
Samples collected before enrichment was enabled keep NULL values in the enrichment columns.

View what a node has been doing in the past minute

Scenario: A node is experiencing high resource utilization, but it’s unclear what kinds of work are running on it and which activities that work is spending its time on. You can query the node-level ASH view to see which work events the node’s workloads have been spending time on:
The query returns the count of samples for each combination of workload type and work event:
The results show this node’s activity is dominated by SQL statements waiting on RPCs to other nodes (DistSenderRemote) and waiting for Raft proposals to be applied (RaftProposalWait), which is typical for write-heavy workloads that must replicate data across nodes. The upsert samples show time spent in the DistSQL processor executing upsert statements, while the ReplicationFlowControl samples indicate that writes are being throttled by . The LockWait samples indicate some on hot keys. To identify which specific statements, jobs, or system tasks are responsible, add workload_id and app_name to the query and group by them.

View cluster-wide workload data from the past 10 minutes

Scenario: Overall cluster resource consumption is high, but it’s unclear which workloads (user queries, background jobs, or system tasks) are responsible for the activity. You can query the cluster-wide ASH view to identify the top workloads consuming resources across all nodes:
The query returns the top 10 workloads by sample count:
The results show that a single SQL statement fingerprint (9bef06d795045524) from the kv application is the largest consumer of cluster resources. System tasks like INTENT_RESOLUTION (async cleanup of transaction intents) and TXN_HEARTBEAT are also significant. To investigate the statement, use the workload_id to find the statement on the . To investigate the job, use its workload_id on the . CockroachDB Cloud users can use the and in the Cloud Console.

Investigate a past incident with persisted ASH data

Scenario: A latency spike occurred yesterday. The in-memory samples from that window have been overwritten, but persisted samples are retained for the retention period. You can query the persisted ASH view for the incident window:
The query returns the top workloads and the work events they were spending time on during the window. To investigate further, identify the dominant work_event values, then use the workload_id to find the statement or job as in the preceding examples.

Find recent lock contention hotspots

Scenario: Elevated p99 latency and increased indicate contention, but it’s unclear which specific workloads are experiencing lock waits and what type of contention is occurring. You can filter ASH samples to show only the lock and latch wait events:
The query shows which workloads experienced lock contention and what type of lock events they encountered:
The results identify the statement fingerprint experiencing latch waits. Use the workload_id to locate the query on the and examine its execution plan and contention time. Review the for contention insights on this statement. CockroachDB Cloud users can use the and in the Cloud Console. If multiple workloads show LockWait events, investigate whether they’re accessing the same tables or rows by examining their query patterns. For detailed contention analysis, see .

Get details about what a specific job is spending time on

Scenario: A background job (such as a , schema change, or ) is running longer than expected, but it’s unclear whether the job is actively doing work or is blocked waiting on something such as locks, admission control, or remote nodes. You can filter by workload type and job ID to understand where the job is spending its time:
The query breaks down the job’s samples by the specific activity it was performing:
The results show the job spent its time in the backup processor (backupDataProcessor) and in replica-level batch evaluation (ReplicaSend), rather than waiting on locks, admission control queues, or remote RPCs. This indicates the job is actively doing work rather than being blocked. To confirm which hardware resources that work is consuming, correlate these results with and for the same time period. To find the job ID for a running job, query the or use : SELECT job_id, description, status FROM [SHOW JOBS]. CockroachDB Cloud users can use the in the Cloud Console.

Known limitations

  • On Basic and Standard CockroachDB Cloud clusters, ASH samples only cover work running on the pod. KV-level work ( I/O, , , etc.) is not visible in ASH samples.

See also