Flagship project / AI + infrastructure

AI with access.
Humans with judgment.

I use AI assistance to administer Gaia, my multi-node Proxmox homelab. The interesting engineering problem isn't simply generating commands. It's defining permission boundaries, recognizing failures across system layers, and retaining a reliable way back.

Proxmox VE 8 → 9Proxmox Backup ServerLinux / LXCNVIDIA / OllamaScoped agent access

01 / Context

The environment

Gaia combines compute nodes Sierra, K2, Fuji and Olympus. The lab spans containerized services, storage, network access, authentication, monitoring and local AI. Olympus provides the NVIDIA GPU capacity used by Ollama in a Linux container.

Human directionIntent · access boundaries · validation
AI assistanceDiagnostics · proposed steps · troubleshooting
Proxmox clusterSierra · K2 · Fuji · Olympus
Workloads and safeguardsContainers · Ollama · backups · monitoring

Conceptual architecture. This does not represent live infrastructure telemetry or publish the underlying access configuration.

02 / Migration journey

Change is a sequence of decisions.

  1. 01

    Define the boundary

    Set human-directed objectives and scoped permissions across the Proxmox servers. An agent recommendation is not the same as authorization to make unrestricted changes.

  2. 02

    Plan for recovery

    Build the backup and recovery approach before the major upgrade, including Proxmox Backup Server and a plan for portable NanoKVM console access.

  3. 03

    Upgrade the cluster

    Use AI assistance to navigate the Proxmox VE 8 to 9 migration, iterating through compatibility concerns and verifying services throughout.

  4. 04

    Restore the workload

    Preserve NVIDIA GPU access through host, device and LXC layers so that Ollama can continue using accelerated local inference.

  5. 05

    Validate the result

    Check the actual post-change behavior rather than treating successful commands or device visibility alone as proof that dependent services work.

03 / The hard parts

GPU access is a chain, not a checkbox.

The Ollama workload depended on compatible behavior across the Proxmox host, NVIDIA components, container access and the application itself. Troubleshooting needed to distinguish a detected device from actual usable model acceleration.

Observed engineering challenge

Maintaining local GPU inference through a hypervisor upgrade and related container/driver changes required iterative investigation across multiple layers.

Operating principle

Validate each dependency at its own boundary, then verify the end-to-end workload. A command that exits successfully is evidence of one step, not proof of the overall outcome.

04 / Recovery discipline

Build a way back.

Proxmox Backup Server was established as part of the recovery approach. A NanoKVM USB was planned as a portable crash cart for direct console access when the network administration path is unavailable; this is not a claim that a full disaster-recovery drill has been completed.

AI can accelerate investigation and execution. Ownership of scope, risk, verification and recovery remains with the operator.

05 / What I learned

Orchestration is engineering work.

Useful AI-assisted administration requires careful system context, limited permissions, explicit acceptance criteria and the willingness to interrogate an apparent success. Those habits transfer directly to technical product work: define intent, understand dependencies and make outcomes observable.

← All technical projects