> For the complete documentation index, see [llms.txt](https://documentation.opencrvs.org/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://documentation.opencrvs.org/v2.1/technical/guides/monitoring.md).

# Monitoring

### 1. Introduction

{% hint style="info" %}
All key OpenCRVS components support the OpenTelemetry tracing standard starting from version 2.1, enabling seamless integration with cloud providers such as AWS, Azure, and GCP, as well as managed monitoring solutions such as Datadog and New Relic.
{% endhint %}

{% hint style="info" %}
**Server-hosted environments only** — These tools are only available for server-hosted environments and are not part of the development environment.
{% endhint %}

Monitoring helps keep an OpenCRVS installation healthy, reliable, and performant by providing visibility into infrastructure, application performance, and system behavior.

OpenCRVS includes pre-installed tools for monitoring and debugging live installations. The [Elastic Stack](https://www.elastic.co/elastic-stack) collects metrics, logs, and alerts, while [Kibana](https://www.elastic.co/kibana) provides access to this observability data.

By default, application logs and infrastructure metrics are retained for the last 30 days, and APM data (traces, errors and service metrics) for the last 7 days. Log and metric retention can be changed with `monitoring.logs.retention_days` and `monitoring.metrics.retention_days` in the dependencies helm chart values, see [Elasticsearch disk management](/v2.1/technical/guides/installation/advanced-topics/elasticsearch-disk-management.md).

This page gives you brief overview how to use preconfigured monitoring solution.

Additional reading:

* For more information how to setup monitoring please visit helm chart documentation [README.md](https://github.com/opencrvs/opencrvs-core/blob/v2.1.0/charts/dependencies/README.md)
* Additional information how to manage disk space for Elasticsearch check [Elasticsearch disk management](/v2.1/technical/guides/installation/advanced-topics/elasticsearch-disk-management.md)

### 2. Monitoring features

OpenCRVS monitoring provides the following core capabilities:

* **Reading and searching application logs** — view detailed logs from all services to debug issues and understand system behaviour.
* **Infrastructure performance insights** — monitor disk space, CPU, and memory usage to know when to scale up.
* **Application Performance Monitoring (APM)** — track service performance, detect bottlenecks, and identify errors.
* **Automated alerting** — receive notifications for application errors and infrastructure health issues.
* **Request tracing** — follow requests through multiple services to understand the full lifecycle.

***

### 3. Getting started with Kibana

Once the environment is installed, the monitoring suite can be accessed using the `kibana.<your_domain>` URL.

#### Login credentials

The login credentials from [GitHub environment creation process](/v2.1/technical/guides/installation/deploy-set-up-a-server-hosted-environment/create-a-github-environment.md):

* `KIBANA_USERNAME` and `KIBANA_PASSWORD`
* username "elastic" and `ELASTICSEARCH_SUPERUSER_PASSWORD`.

### 4. Monitoring tools

OpenCRVS uses several specialized tools as part of the monitoring stack:

#### 4.1 Metricbeat

Metricbeat gets installed on all host machines in your infrastructure. Its sole purpose is to collect data about the network, the host machines, and the Kubernetes environment. The data is stored in the OpenCRVS Elasticsearch database.

This data can be viewed by navigating to **Observability** → **Metrics** and selecting either **Inventory** or **Metrics Explorer**. The data can be visualized, grouped, and filtered in these views.

#### 4.2 Application Performance Monitoring (APM)

The OpenCRVS monitoring stack comes with a pre-installed Application Performance Monitoring tool (APM). This tool collects performance metrics, errors, and HTTP request information from each of the services in the OpenCRVS stack.

You can find this tool in Kibana by navigating to **Observability** → **APM** → **Services**. This tool can be used to:

* Catch anomalies such as errors happening inside the services
* Detect bottlenecks in the architecture
* Identify which services should be scaled up

#### 4.3 How traces are collected (OpenTelemetry)

OpenCRVS core services, the NGINX servers of the client and login applications, and the Traefik ingress controller are instrumented with [OpenTelemetry](https://opentelemetry.io/). They send traces and metrics over OTLP (gRPC) to an OpenTelemetry Collector running in the dependencies namespace. The collector forwards the data to the Elastic APM Server, which stores it in Elasticsearch, where it appears under **Observability** → **APM** in Kibana:

```
OpenCRVS services / NGINX / Traefik
        │  OTLP gRPC (port 4317)
        ▼
opentelemetry-collector.opencrvs-deps-<environment>.svc.cluster.local
        │
        ▼
apm-server (port 8200) → Elasticsearch → Kibana APM
```

The collector is installed by the **Deploy dependencies** GitHub Actions workflow of your infrastructure repository when `environments/<environment>/opentelemetry/values.yaml` exists. This file is generated by the `yarn environment:init` script.

Tracing is configured in the `otel` section of `environments/<environment>/opencrvs-services/values.yaml`:

```yaml
otel:
  enabled: true
  deployment_environment: <environment>
  exporter_otlp_endpoint: "opentelemetry-collector.opencrvs-deps-<environment>.svc.cluster.local:4317"
```

* `enabled`: Enables OpenTelemetry instrumentation. `yarn environment:init` sets it to `true`.
* `deployment_environment`: Environment name reported with every trace. Defaults to `NODE_ENV`.
* `exporter_otlp_endpoint`: Collector address. Required when `enabled` is `true`.
* `exporter_otlp_protocol`: (Optional) Defaults to `grpc`.

Only traces and metrics are sent through OpenTelemetry. Application logs are still collected by Filebeat.

{% hint style="info" %}
To send traces to another backend, such as a cloud provider or a managed monitoring service (Datadog, New Relic), add an exporter to the collector configuration in `environments/<environment>/opentelemetry/values.override.yaml` and redeploy dependencies.
{% endhint %}

### 5. Monitoring topics

The following pages provide detailed guidance on specific monitoring topics:

#### [Application logs](/v2.1/technical/guides/monitoring/application-logs.md)

Learn how to access, search, and trace application logs to debug issues and understand system behavior.

#### [Infrastructure health](/v2.1/technical/guides/monitoring/infrastructure-health.md)

Monitor critical infrastructure metrics such as disk space, CPU usage, and memory consumption to proactively manage resources.

#### [Routine monitoring](/v2.1/technical/guides/monitoring/routine-monitoring.md)

Establish daily monitoring practices and understand the built-in alerts to maintain a healthy installation.

#### [Setting up alerts](/v2.1/technical/guides/monitoring/setting-up-alerts.md)

Configure custom alerts to notify you when critical conditions occur, ensuring rapid response to issues.

***

### 6. Read more

* [Elasticsearch disk management](/v2.1/technical/guides/installation/advanced-topics/elasticsearch-disk-management.md)
* [OpenCRVS Dependencies and monitoring helm chart README.md](https://github.com/opencrvs/opencrvs-core/blob/v2.1.0/charts/dependencies/README.md)
* [Kibana — your window into Elastic](https://www.elastic.co/guide/en/kibana/current/introduction.html#introduction)
* [Application Performance Monitoring (APM)](https://www.elastic.co/observability/application-performance-monitoring)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://documentation.opencrvs.org/v2.1/technical/guides/monitoring.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
