Remote observability on linux devices #9

Open
opened 2023-10-06 11:22:12 +00:00 by janhalen · 5 comments
janhalen commented 2023-10-06 11:22:12 +00:00 (Migrated from github.com)

👤 User story

As the administrator i need to be able to Monitor and detect errors and/or misconfigurations in remote devices early, To be able to meet the expectations, that devices have a high uptime and availability and that the devices are consistently compliant with the given security policies.

📋 Tasks and features

  • research and prototype a standardized graphical dashboarding tool
    - [ ] Take a look at OpenObserve
  • select a standards based metrics/log collection method - @janhalen: OTEL is the obvious choice from my POV
  • add a notification system
  • Define a minimal set of metrics for devices
### 👤 User story <p>As the administrator i need to be able to Monitor and detect errors and/or misconfigurations in remote devices early, To be able to meet the expectations, that devices have a high uptime and availability and that the devices are consistently compliant with the given security policies.</p> ### 📋 Tasks and features - [ ] research and prototype a standardized graphical dashboarding tool - [ ] Take a look at [OpenObserve](https://github.com/openobserve/openobserve) - [ ] select a standards based metrics/log collection method - @janhalen: [OTEL](https://opentelemetry.io/) is the obvious choice from my POV - [ ] add a notification system - [ ] Define a minimal set of metrics for devices
janhalen commented 2023-10-06 12:37:25 +00:00 (Migrated from github.com)

pre-development branch experiments are done in a codespace here-> https://glorious-pancake-wgpxq79g54pcxrr.github.dev/

Findings and results of pre-dev prototype will be added to a dev-branch here as soon as all pending PRs are merged in.

pre-development branch experiments are done in a codespace here-> https://glorious-pancake-wgpxq79g54pcxrr.github.dev/ Findings and results of pre-dev prototype will be added to a dev-branch here as soon as all pending PRs are merged in.
janhalen commented 2026-01-12 07:45:40 +00:00 (Migrated from github.com)

Status Q1 2026:

From a 12/15 factor design principle and utilizing a Cloud Native approach @turegjorup and I agreed that integration with standardized Observabilty components like OTEL would be the right way forward.

This will shift much of the work needed from core logging and observability code, to standardised integration configurations.

Status Q1 2026: From a 12/15 factor design principle and utilizing a Cloud Native approach @turegjorup and I agreed that integration with standardized Observabilty components like [OTEL](https://opentelemetry.io/) would be the right way forward. This will shift much of the work needed from core logging and observability code, to standardised integration configurations.
janhalen commented 2026-02-20 11:33:28 +00:00 (Migrated from github.com)

@ChatBotBerg: Taking a look at OpenObserve when i get the time.

@ChatBotBerg: Taking a look at OpenObserve when i get the time.
rriemann commented 2026-03-19 20:22:58 +00:00 (Migrated from github.com)

Hi all,

this is an important point. For the PoC described at https://eu-os.eu/poc/ we rely on Foreman, as it solves many concerns at once and supports already bootc.

https://eu-os.eu/poc/manage-fleet/

For our PoC we demonstrated and documented already:

  • PXE provisioning of bootc images built in Gitlab to test devices (Thinkpad laptops)
  • during this provisioning, we enrol the laptop both in FreeIPA and in Foreman. Foreman enrolment relies on subscription-manager (FOSS version of redhat subscription manager).
  • subcription-manager and the systemd job (comes with the bootc RPM) bootc-publish-rhsm-facts.service send fact files on each boot (configurable) to Foreman. This gives observability.
  • Foreman supports running custom tasks on hosts and can collect feedback. See: https://theforeman.org/2021/02/introduction-to-the-remote-execution-plugin.html

What is missing in our setup is:

  • how can we conveniently apply configuration to individual or sets of laptops?
  • how can we already setup upon provisioning a Wireguard VPN network and then use Foreman cockpit to get a web shell into laptops for trouble shooting?
Hi all, this is an important point. For the PoC described at https://eu-os.eu/poc/ we rely on Foreman, as it solves many concerns at once and supports already bootc. https://eu-os.eu/poc/manage-fleet/ For our PoC we demonstrated and documented already: - PXE provisioning of bootc images built in Gitlab to test devices (Thinkpad laptops) - during this provisioning, we enrol the laptop both in FreeIPA and in Foreman. Foreman enrolment relies on subscription-manager (FOSS version of redhat subscription manager). - subcription-manager and the systemd job (comes with the bootc RPM) `bootc-publish-rhsm-facts.service` send fact files on each boot (configurable) to Foreman. This gives observability. - Foreman supports running custom tasks on hosts and can collect feedback. See: https://theforeman.org/2021/02/introduction-to-the-remote-execution-plugin.html What is missing in our setup is: - how can we conveniently apply configuration to individual or sets of laptops? - how can we already setup upon provisioning a Wireguard VPN network and then use Foreman cockpit to get a web shell into laptops for trouble shooting?
janhalen commented 2026-03-25 10:09:37 +00:00 (Migrated from github.com)

Hi @rriemann!

As you probably have figured by now the reference architecture for this poc project is anchored in Cloud Native Operational principles, namely GitOps. The major benefits of operating a device fleet with GitOps are mainly:

  • Single Source of Truth: Uses Git as the definitive state, eliminating configuration drift across the fleet.
  • Auditability: Provides a complete historical record of every change for simplified compliance and security.
  • Rapid Recovery: Enables near-instant rollbacks to known stable states via standard Git operations.
  • Declarative Scaling: Manages desired states rather than individual units, allowing small teams to oversee massive fleets.
  • Pull-Based Security: Devices pull updates from Git, removing the need for open inbound ports and reducing the attack surface.
Hi @rriemann! As you probably have figured by now the reference architecture for this poc project is anchored in Cloud Native Operational principles, namely GitOps. The major benefits of operating a device fleet with GitOps are mainly: * **Single Source of Truth:** Uses Git as the definitive state, eliminating configuration drift across the fleet. * **Auditability:** Provides a complete historical record of every change for simplified compliance and security. * **Rapid Recovery:** Enables near-instant rollbacks to known stable states via standard Git operations. * **Declarative Scaling:** Manages desired states rather than individual units, allowing small teams to oversee massive fleets. * **Pull-Based Security:** Devices pull updates from Git, removing the need for open inbound ports and reducing the attack surface.
Sign in to join this conversation.
No description provided.