Skip to main content
This page covers installing the agent in networks that restrict outbound traffic: an HTTP proxy is mandatory, a TLS-intercepting proxy re-signs traffic with a corporate CA, internal services use a private CA, or the host cannot pull from public registries.

Outbound requirements

The agent only makes outbound connections. Nothing needs to reach the agent from the internet, and no inbound firewall rule is required. The agent does not follow redirects from Observer Cloud, so a proxy or gateway that answers with a 3xx (for example a captive portal or a login page) shows up as a failed request in the agent logs, not as a silent success. Everything the agent sends to the cloud is described in The agent / cloud boundary.

Use an HTTP proxy

The agent’s runtime honours the standard proxy variables. Set them in the agent’s environment: If the proxy needs credentials, put them in the URL: http://user:[email protected]:3128. Treat the value as a secret.

Keep internal probe targets out of the proxy

The proxy variables apply to every HTTP request the agent makes, not just the connection to Observer Cloud. That includes Prometheus queries and HTTP probes. Most probe targets are internal services that the proxy either cannot reach or should not see, so add them to NO_PROXY:
  • The Prometheus host from PROMETHEUS_SERVER_URL.
  • The hosts of internal HTTP, Loki and Elasticsearch targets.
  • localhost and 127.0.0.1.
  • In Kubernetes, the cluster domain suffix (for example .svc,.svc.cluster.local), so in-cluster Services are reached directly.
Probes that open their own connections (TCP, DNS, ICMP, TLS certificate and database probes) connect directly to the target and are not routed through the proxy. gRPC and WebSocket probes use their own client libraries; confirm their behaviour with a test probe before relying on the proxy for them.

TLS-intercepting proxies and private CAs

A TLS-intercepting proxy (also called TLS inspection or SSL inspection) terminates HTTPS and re-signs it with a corporate CA. The agent rejects that certificate unless it trusts the CA, and logs a certificate verification error on every request to the cloud. Add the CA to the agent’s trust store with NODE_EXTRA_CA_CERTS: point it at a PEM file on the agent host (or inside the container) that holds the CA certificate, or several certificates concatenated. The runtime trusts these certificates in addition to its built-in trust store, so public endpoints keep working. The same variable covers internal HTTPS targets signed by a private CA: Prometheus, Loki, Elasticsearch and HTTP probes all trust the extra certificates. NODE_EXTRA_CA_CERTS is read when the process starts. After you change the file or the variable, restart the agent. Ask your network or security team for the CA certificate in PEM format (a text file that starts with -----BEGIN CERTIFICATE-----). If you only have a DER file (.cer or .crt with binary content), convert it:
The container image runs as a non-root user (UID 65532), so make sure a mounted PEM file is world-readable (chmod 644).

Per-probe CA certificates

Some probes take a CA certificate in their own configuration instead. This scopes the trust to that one probe:
  • gRPC probes: set ca_cert_ref in TLS or mTLS mode. See gRPC probes.
  • HTTP probes: ca_cert_ref applies together with mTLS client certificates. See mTLS authentication. For a plain HTTPS target behind a private CA, use NODE_EXTRA_CA_CERTS.
ca_cert_ref holds the name of an environment variable on the agent host, never the certificate itself. The variable’s value is either the PEM text or a path to a PEM file.

Do not disable TLS verification

Two variables turn certificate checks off. Neither belongs in production:
  • SKIP_SSL_VERIFICATION=true disables verification on the connection to Observer Cloud. Anyone who can intercept that connection can read your agent key and feed the agent false metric definitions.
  • NODE_TLS_REJECT_UNAUTHORIZED=0 disables verification for every HTTPS request the agent makes: the cloud, Prometheus, and every probe. An HTTP probe then reports a target as healthy even when its certificate is invalid or forged.
If verification fails, fix the trust store with NODE_EXTRA_CA_CERTS instead. To accept a self-signed certificate on a single internal HTTP probe, set verify_tls: false on that probe only.

Examples

Each example sets a proxy, bypasses it for internal targets, and trusts a corporate CA stored in corp-ca.pem. Remove the parts you do not need.

Install without access to public registries

The host that runs the agent does not need to reach ghcr.io or github.com after installation. Fetch the artifact on a machine that can, then move it across.

Container image: save and load

On a machine with internet access, pull the image for the architecture of the target host and save it to a file:
Use --platform linux/arm64 for ARM hosts. Copy the tar file to the target host, check the checksum, and load it:
Then run the agent with the same image reference.

Container image: mirror into an internal registry

If your organisation runs its own registry, copy the image into it and point the install at the mirror. ghcr.io/useobserver/agent is a multi-architecture image (linux/amd64 and linux/arm64); use a tool that copies every architecture, such as skopeo or crane:
Replace ghcr.io/useobserver/agent with registry.corp.example/observer/agent in the Docker, Compose or Kubernetes examples. Pin exact versions so every host runs the image you mirrored.

Release binary: download and transfer

On a machine with internet access, download the binary and the checksum file from github.com/useobserver/agent/releases:
Copy both files to the target host and verify the binary there, after the transfer:
Then continue with the systemd steps in Install from a release binary.

Air-gapped and isolated networks

Observer Cloud is a hosted service, so the agent needs an outbound HTTPS path to it, directly or through a proxy. No inbound connection is ever required. A network with no route to Observer Cloud at all is not supported: the agent can run and collect readings there, but nothing reaches your status pages. What the agent does tolerate is losing that path for a while. Every result is written to the local queue first, a SQLite file at BUFFER_PATH, and a background process delivers it to the cloud in batches. When the cloud is unreachable:
  • The agent keeps probing and keeps queueing results.
  • Delivery retries with exponential backoff, from 1 second up to a maximum of 5 minutes between attempts.
  • The queue holds up to BUFFER_MAX_ROWS rows (default 10000). Once it is full, the oldest rows are evicted to make room and the agent logs an error with the number dropped.
  • So that one bad row cannot block the queue forever, a row that fails 20 delivery attempts in a row is dropped with an error log. During a long outage this costs at most one row roughly every hour or two.
  • When the path returns, the backlog drains in batches of up to 100 rows per request, oldest first.
How long an outage the default cap covers depends on how many probes the agent runs and how often. Raise BUFFER_MAX_ROWS if you expect long interruptions, and keep BUFFER_PATH on a persistent volume so the queue survives a restart. While the agent cannot reach the cloud it also stops sending heartbeats, so the console marks it offline and the agent.offline alert fires even though probing continues locally. For jobs on hosts that cannot reach the internet at all, the heartbeat relay lets them report heartbeat checks through the agent instead.

Verify connectivity

After starting the agent, check three places. Agent logs. A working connection logs Successfully fetched metric definitions within about 30 seconds. A failed one logs Error fetching metric definitions: followed by the cause. Heartbeat failures are only logged with VERBOSE=true, which is worth setting while you debug the network path. Typical causes: Where to read the logs: docker logs -f observer-agent, docker compose logs -f observer-agent, kubectl -n observer logs -f deploy/observer-agent, or journalctl -u observer-agent -f. Local dashboard. The agent serves a read-only dashboard on port 10101, bound to loopback by default. Its Cloud link card shows the last heartbeat and the last push, with the error text when one failed, and the queue card shows whether results are piling up. See Read the agent dashboard. Console. The agent’s page in the console shows when it last checked in. A connected agent is marked running within 90 seconds of starting. If it stays offline while the logs show no errors, see Diagnose a stalled agent. To test the network path from the agent host before installing, make a request through the same proxy. Pass --cacert only when a TLS-intercepting proxy re-signs the traffic, because it replaces curl’s default trust store:
Any HTTP status code means the path and the certificate chain work. A TLS error means the CA file is wrong or incomplete.