// Platform & Site Reliability Engineering
Brady Bangasser
I build and operate large-scale systems to stay fast, available, and secure, across HPC, multi-cloud infrastructure, and the automation that ships them.

// focus areas
Reliability & Platform
High-availability systems, observability, and self-healing infrastructure built to survive failure — not just avoid it.
Cloud & Infrastructure
Multi-cloud and on-prem infrastructure as code, least-privilege IAM, and security hardened to recognized standards.
HPC & Distributed Systems
Cluster operations at scale — SLURM and Kubernetes scheduling, performance tuning, and distributed workloads.
Systems & Security
Low-level performance work, compiler and toolchain internals, and applied cryptography and authentication.
// writing
view all →Building a rootless edge: Traefik, Podman, and per-container TLS
How I built a self-registering reverse proxy on rootless Podman with per-domain TLS, a private CA, and multi-cloud portability, and where I chose it over Kubernetes.
Notes on job scheduling in HPC clusters
A deeper look at how scheduling policy shapes throughput and fairness on shared clusters.
// consulting