Troubleshooting sidecar performance issues with pprof

When requested by Splunk Support, use Go built-in pprof tool to collect troubleshooting data for performance issues with the supervisor and supported sidecars.

Use the Go pprof tool to collect troubleshooting data for performance issues with the supervisor and supported sidecars. The tool captures runtime performance metrics, called profiles, that include CPU usage, memory allocations, goroutine activity, and mutex contention. This data helps Splunk developers identify performance bottlenecks and optimize resource consumption.

In the Splunk platform, the sidecar runs an internal HTTP server that exposes pprof endpoints. A Splunk Support engineer might ask you to use these endpoints to collect profiles during an incident.

The pprof tool is active by default for the supervisor and some sidecars, such as Agent Management, Data Orchestration (DO), IPC Broker, and Spotlight. As a result, you can collect diagnostic data during active incidents without restarting processes. This capability improves incident response and troubleshooting efficiency.

Configuring pprof through Unix domain sockets

Sidecars can expose pprof servers through Unix domain sockets for local-only access. The system automatically creates socket files when a sidecar starts and removes them when it stops.

Sidecars expose pprof servers through a network port or a Unix domain socket (socket file). When using a socket file, the profiling server listens locally on this file instead of a TCP port.

Using a Unix domain socket restricts access to the local host, ensuring that only users with the appropriate operating system permissions can interact with the pprof server. The owner of the domain socket is the user running the Splunk platform process.

Socket file location

The socket files are located at the following path: $SPLUNK_HOME/var/run/supervisor/pprof-<component-name>.sock.

The <component-name> variable can be supervisor or the sidecar name used in the pprof configuration. Component names may differ from sidecar process names. For process names, see Sidecar configuration settings Sidecar configuration settings.

File location path - example Notes
$SPLUNK_HOME/var/run/supervisor/pprof-supervisor.sock

The component name is supervisor.

$SPLUNK_HOME/var/run/supervisor/pprof-agent-manager.sock

The component name of the Agent Management sidecar is agent-manager.

Socket files - lifecycle

The system creates the socket files automatically when a sidecar starts and removes them when it stops. If you deactivate the pprof tool for a component (the supervisor or a sidecar), the system does not create a socket file.

Profiles available through pprof

View the runtime profiles that the pprof tool exposes, the data each profile captures, and the endpoints used to collect them.

A runtime profile is a collection of data that captures the behavior and performance characteristics of a process while it is running. The following profiles are available through the pprof tool:

Profile name Description and interpretation Endpoint
Allocs

A sample of all past heap allocations.

A high value of allocation churn means the application is frequently creating and deleting objects, which increases the CPU and system resources required by the Garbage Collector (GC).

http://localhost/debug/pprof/allocs
Block

Time spent blocked on synchronization operations, such as channels, locks, select statements, or timers.

A high block value means that synchronization logic is causing significant latency.

http://localhost/debug/pprof/block
Cmdline

The command line arguments used to start the current process.

A value of unexpected command-line arguments means the process may have been started with incorrect configuration or environment variables.

http://localhost/debug/pprof/cmdline
CPU

Functions consuming the most CPU time

http://localhost/debug/pprof/profile
Heap

Live objects currently allocated on the heap.

A value of increasing heap size over time means the application may have a memory leak.

http://localhost/debug/pprof/heap
Goroutine

Stack traces of all current goroutines

Helps in understanding what each goroutine is doing, which functions are active, and where in the code execution is occurring.

http://localhost/debug/pprof/goroutine
Mutex

Mutexes (mutual exclusions) currently contended and the amount of delay they introduce.

A high mutex value means multiple goroutines are often waiting for the same lock.

http://localhost/debug/pprof/mutex
Symbol

Mapping of program counter addresses in the compiled code to function names for human-readable profile analysis.

With clearly mapped function names, you can quickly identify the specific code causing a performance issue.

http://localhost/debug/pprof/symbol
Threadcreate

Stack traces that led to the creation of new OS threads (threads managed and scheduled by the operating system).

A high value of thread creation rate means the application is creating too many background threads, which can slow down the system and use more memory than necessary.

http://localhost/debug/pprof/threadcreate
Trace

A time-ordered execution trace, collecting scheduling, system calls, and garbage collection activity.

Long pauses in the execution timeline mean the application is stuck waiting for tasks to start or finish, which causes slow performance.

http://localhost/debug/pprof/trace

Collect a runtime profile through the pprof tool

During an incident, a Splunk Support engineer might ask you to use the pprof tool to collect runtime profiles from the supervisor and sidecars.

During an incident, a Splunk Support engineer might ask you to collect profiles directly from a component by using curl with the --unix-socket option. The command saves the binary output to your current working directory using the naming convention <profile-name>-profile.out.

In the following commands, replace component-name with the name of the sidecar specified by the Splunk Support engineer.

To retrieve the CPU profile, run the following command:

Note: By default, the CPU profile runs for 30 seconds.
CODE
curl --unix-socket $SPLUNK_HOME/var/run/supervisor/pprof-<component-name>.sock \
  http://localhost/debug/pprof/profile --output ./cpu-profile.out
To retrieve the Heap profile, run the the following command:
CODE
curl --unix-socket $SPLUNK_HOME/var/run/supervisor/pprof-<component-name>.sock \
  http://localhost/debug/pprof/heap --output ./heap-profile.out

To retrieve the Goroutine profile, run the the following command:

CODE
curl --unix-socket $SPLUNK_HOME/var/run/supervisor/pprof-<component-name>.sock \
  http://localhost/debug/pprof/goroutine --output ./goroutine-profile.out

Troubleshoot errors with too long socket paths

If you encounter a curl (6) Unix socket path too long error, the path to the socket file exceeds the system limit.
CODE
curl --unix-socket <long-path>/pprof-<component-name>.sock \
  http://localhost/debug/pprof/heap --output ./heap-profile.out
curl: (6) Unix socket path too long: '<long-path>/pprof-<component-name>.sock'

component-name is the value of supervisor or the sidecar name used in the pprof configuration.

You can resolve this by creating a shorter symbolic link (symlink) that points to the actual socket file:

  1. Create a link in a temporary directory, for example /tmp that points to the socket file:
    CODE
    ln -s <long-path>/pprof-<component-name>.sock /tmp/<component-name>.sock
    where:
    • <long-path>/pprof-<component-name>.sock is the socket file path.

    • /tmp/<component-name>.sock is the shorter link name.

  2. Collect the profile using the symlink:

    CODE
    curl --unix-socket /tmp/<component-name>.sock \
      http://localhost/debug/pprof/heap --output ./heap-profile.out

Configuration settings of the pprof tool

Configure global and per-component pprof settings in the server.conf file.

You can configure pprof settings in server.conf using the following stanzas:
  • [sidecars_pprof] - Specifies global settings for the supervisor and all sidecars.

  • [sidecars_pprof:<component_name>] - Specifies per-component settings (for the supervisor or a specific sidecar).

    Note: Per-component settings override global settings. Global settings override the default ones.

Configuration settings

The following table presents pprof settings that you can specify.

Setting Type Default value Description
enabled boolean true

A value of true starts the pprof server and activates endpoints.

A value of false means that the pprof server does not start, no socket file is created and no endpoints are available.

mutex_fraction integer 0

Use this setting to activate mutex contention profiling. It shows how often processes compete for the same mutex lock, causing delays.

A value greater than 0 activates sampling at a rate of 1/mutex_fraction.

A value of 0 deactivates mutex profiling.

Note: Activating this setting may impact performance.
block_rate integer 0

Use this setting to activate goroutine block profiling.

A value of 0 deactivates block profiling.

A value of 1 captures every blocking event.

A value greater than 1 captures blocking events at the rate of 1/block rate of nanoseconds.

Note: Activating this setting may impact performance.

Example configuration

The following settings activate mutex and block profiling globally and deactivate pprof for the supervisor:
CODE
# Global settings for all components
[sidecars_pprof]
mutex_fraction = 1000
block_rate = 1000000

# Override: Deactivate pprof for the supervisor
[sidecars_pprof:splunk-supervisor]
enabled = false

Deactivate the pprof tool

Deactivate the pprof tool globally or per component in server.conf.

You can deactivate the pprof tool globally or for specific components (the supervisor or a specific sidecar) in server.conf.

To deactivate pprof for the supervisor and all sidecars, specify the following in the global stanza:

CODE
[sidecars_pprof]
enabled = false
To deactivate the pprof tool for a component, specify the following in the component-specific stanza:
CODE
[sidecars_pprof:<component_name>]
enabled = false

For component_name, use supervisor or the name of the sidecar that the Splunk Support engineer asks you to investigate.

When pprof is deactivated, the component does not initialize the web server, no socket file is created, and any additional pprof settings in that stanza are ignored.