- Infrastructure metric
- 70+
B200/B300-class systems installed
AI Infrastructure & GPU Systems
AI Infrastructure & Systems Engineer
Building and operating AI infrastructure and GPU systems across Linux, Ceph and distributed storage, virtualization, high-speed networking, data-center deployment, security, telemetry, and observability—while developing bare-metal automation and platform capabilities.
Salt Lake City, Utah
Selected evidence
Approved aggregate facts from professional infrastructure work. Customer, facility, topology, and system-identifying details remain excluded.
B200/B300-class systems installed
GPU systems supported across mixed NVIDIA environments
Completed professional Ceph environment
Approximate raw Ceph capacity
AI/GPU infrastructure first
The hierarchy starts with AI and GPU systems, then Linux, storage, networking, security and observability; automation and platform capabilities remain active development areas.
Selected work
In-progress lab development sits alongside completed and ongoing professional work. Public case studies use aggregate evidence and abstract confidential operating context.
An abstracted case study of ongoing professional GPU infrastructure work: more than 70 B200/B300-class systems installed, hundreds of GPU systems supported, and experience across DGX, HGX, GB200, H200, B200, and B300-class environments.
An in-progress lab expansion and design effort for developing bare-metal provisioning, out-of-band management, Kubernetes foundations, telemetry, and operational runbooks without overstating deployed inventory.
An abstracted case study of completed professional work with an eight-node Ceph environment providing approximately 100–200 TB of raw capacity for VMware storage, backups, historical archive, and retention needs.
An abstracted case study of completed professional security and observability work spanning Nessus, PKI/TLS, CMRS-related work, secure data-diode collection, Grafana, Prometheus, IPMI, SNMP, and legacy BMS/BAS telemetry integration.
Experience
A concise view of current and recent infrastructure roles. The full experience page separates operational environment, responsibilities, evidence, and confidentiality boundaries.
Feb 2026 – Present
Operates AI and hyperscale data-center infrastructure across GPU compute, high-speed interconnects, power, cooling, telemetry, and incident response.
Current role
Oct 2023 – Jun 2025
Led infrastructure security, systems engineering, and operational support across federal cloud, virtualization, distributed storage, monitoring, and legacy environments.
Mar 2022 – Oct 2023
Supported high-availability financial data-center infrastructure under strict uptime, change-control, physical-security, and incident-response requirements.
In progress
These are active learning and proof-of-work directions, not claims of completed mastery.
Active development focus: building fluency in out-of-band management interfaces, lifecycle workflows, and safe automation patterns.
Active development focus: designing repeatable discovery, boot, provisioning, validation, and recovery workflows for Linux hosts.
Active development focus: understanding cluster foundations, accelerator scheduling, device visibility, and operator runbooks.
Active development focus: turning infrastructure intent into reviewable, repeatable configuration and change workflows.
Active development focus: developing 25/100 Gb design judgment across topology, optics, cabling, validation, and failure isolation.
Active development focus: connecting actionable telemetry to runbooks, controlled response, and post-incident learning.
Knowledge sharing
Only complete, public articles appear here; internal drafts never surface as published work.
An operator’s model for Ceph capacity, failure domains, recovery headroom, health evidence, and safe growth decisions.