AI Infrastructure & GPU Systems

Greyson K. Evers

AI Infrastructure & Systems Engineer

Building and operating AI infrastructure and GPU systems across Linux, Ceph and distributed storage, virtualization, high-speed networking, data-center deployment, security, telemetry, and observability—while developing bare-metal automation and platform capabilities.

Salt Lake City, Utah

Selected evidence

Infrastructure scale

Approved aggregate facts from professional infrastructure work. Customer, facility, topology, and system-identifying details remain excluded.

  • Infrastructure metric
    70+

    B200/B300-class systems installed

  • Infrastructure metric
    Hundreds

    GPU systems supported across mixed NVIDIA environments

  • Infrastructure metric
    8 nodes

    Completed professional Ceph environment

  • Infrastructure metric
    100–200 TB

    Approximate raw Ceph capacity

AI/GPU infrastructure first

Infrastructure expertise

The hierarchy starts with AI and GPU systems, then Linux, storage, networking, security and observability; automation and platform capabilities remain active development areas.

  1. AI systems and GPU platforms

    • GPU server bring-up
    • AI infrastructure operations
    • DGX/HGX-class environment familiarity
    • GPU telemetry and validation
    • AI platform support workflows
  2. Linux and bare-metal infrastructure

    • Linux administration
    • systemd and host troubleshooting
    • bare-metal provisioning concepts
    • BMC/IPMI/Redfish development focus
    • PXE development focus
  3. Storage and virtualization

    • Ceph distributed storage
    • VMware infrastructure
    • block/object/file storage concepts
    • capacity and recovery planning
    • Kubernetes storage learning path
  4. Networking and interconnects

    • high-speed networking
    • MPO/MTP fiber familiarity
    • AOC/DAC familiarity
    • 25/100 Gb network design learning path
    • optical/cabling BOM review experience
  5. Security and observability

    • Nessus lifecycle exposure
    • PKI/TLS and certificate lifecycle
    • identity controls and hardening
    • Prometheus and Grafana
    • IPMI and SNMP telemetry
  6. Automation and platform operations

    • Bash and scripting
    • Python automation development focus
    • Infrastructure-as-code development focus
    • Kubernetes platform development focus
    • repeatable operational runbooks

Selected work

Infrastructure case studies

In-progress lab development sits alongside completed and ongoing professional work. Public case studies use aggregate evidence and abstract confidential operating context.

  1. An abstracted case study of ongoing professional GPU infrastructure work: more than 70 B200/B300-class systems installed, hundreds of GPU systems supported, and experience across DGX, HGX, GB200, H200, B200, and B300-class environments.

    High-Density GPU Deployment and Operational Workflow categories:
    • AI infrastructure
    • GPU systems
    • Data-center operations
  2. An in-progress lab expansion and design effort for developing bare-metal provisioning, out-of-band management, Kubernetes foundations, telemetry, and operational runbooks without overstating deployed inventory.

    AI Infrastructure Homelab and Bare-Metal Control Plane categories:
    • AI infrastructure
    • Bare metal
    • Automation development
  3. An abstracted case study of completed professional work with an eight-node Ceph environment providing approximately 100–200 TB of raw capacity for VMware storage, backups, historical archive, and retention needs.

    Ceph-Backed Production Infrastructure categories:
    • Storage
    • Linux
    • Distributed storage
  4. An abstracted case study of completed professional security and observability work spanning Nessus, PKI/TLS, CMRS-related work, secure data-diode collection, Grafana, Prometheus, IPMI, SNMP, and legacy BMS/BAS telemetry integration.

    Secure Infrastructure Telemetry and Observability categories:
    • Security
    • Observability
    • Infrastructure operations

Experience

Selected professional experience

A concise view of current and recent infrastructure roles. The full experience page separates operational environment, responsibilities, evidence, and confidentiality boundaries.

  1. DataBank

    Senior Data Center Operations Engineer

    Salt Lake City, Utah

    Feb 2026 – Present

    Operates AI and hyperscale data-center infrastructure across GPU compute, high-speed interconnects, power, cooling, telemetry, and incident response.

    Current role

  2. Centers for Disease Control and Prevention

    Infrastructure Security Lead (GS-12)

    Atlanta, Georgia

    Oct 2023 – Jun 2025

    Led infrastructure security, systems engineering, and operational support across federal cloud, virtualization, distributed storage, monitoring, and legacy environments.

  3. BGIS / PayPal data-center environment

    Senior Data Center Engineer

    Salt Lake City, Utah

    Mar 2022 – Oct 2023

    Supported high-availability financial data-center infrastructure under strict uptime, change-control, physical-security, and incident-response requirements.

In progress

Current engineering development

These are active learning and proof-of-work directions, not claims of completed mastery.

  1. BMC, IPMI, and Redfish

    Active development focus: building fluency in out-of-band management interfaces, lifecycle workflows, and safe automation patterns.

  2. PXE and bare-metal provisioning

    Active development focus: designing repeatable discovery, boot, provisioning, validation, and recovery workflows for Linux hosts.

  3. Kubernetes and GPU scheduling

    Active development focus: understanding cluster foundations, accelerator scheduling, device visibility, and operator runbooks.

  4. Infrastructure as code

    Active development focus: turning infrastructure intent into reviewable, repeatable configuration and change workflows.

  5. High-speed networking

    Active development focus: developing 25/100 Gb design judgment across topology, optics, cabling, validation, and failure isolation.

  6. Observability and automated remediation

    Active development focus: connecting actionable telemetry to runbooks, controlled response, and post-incident learning.

Knowledge sharing

Recent technical writing

Only complete, public articles appear here; internal drafts never surface as published work.

Next step

Building AI infrastructure with systems depth

Review selected infrastructure work and the semantic HTML résumé, or reach out directly about AI infrastructure and systems engineering opportunities.