Linux Sysadmin – Debug Random Shutdowns on BRIX Hardware -- 2

Job ID: 40555507

Budget: $2 – $8 USD

We have 13 BRIX mini-PC machines running Debian 12 that are experiencing random, unexplained shutdowns across multiple geographic locations. The remaining 47 servers in our fleet (different hardware) run without any issues. Standard OS-level logs show nothing useful, suggesting the shutdowns are occurring below the OS layer. We need an experienced Linux/infrastructure engineer to identify the root cause and deliver a documented fix.

Requirements:
- Strong experience with Linux (Debian/Ubuntu) at the system and kernel level
- Familiarity with ACPI, power management, and hardware-firmware interaction debugging
- Experience reading BMC/IPMI/SEL hardware event logs
- Ability to audit and interpret kernel ring buffer (journald/dmesg) output
- Experience with Ansible-managed infrastructure
- Comfortable working with physical or remote-access hardware across distributed locations
- Prior experience debugging hardware-specific Linux issues (not just software-level)

Deliverables:
- Root cause analysis report identifying why the BRIX units are shutting down
- Documented fix or remediation steps that can be applied across all 13 units
- Recommendations for monitoring to catch and alert on future events before they cause downtime
- Any Ansible playbook changes needed to prevent recurrence

Work arrangement:
This is an hourly engagement. We expect to start with a single BRIX machine for initial investigation, then expand to the full fleet once root cause is confirmed. Estimated scope is small-to-medium depending on how quickly the issue can be reproduced and traced.

About the project:
This is a live production infrastructure issue affecting 13 machines across multiple sites. The hardware runs the same Debian 12 image deployed via Ansible, and only the BRIX units are affected — making this a hardware-specific debugging challenge that requires someone comfortable working at the firmware and kernel boundary.