Come and join the MDC Team where we thrive on data to solve high-impact business problems!
About the Role:
The candidate will own the path from local development to deployed AWS cluster across DCSA's GovCloud (IL2/IL5) and classified (IL6/Secret) partitions, and the health of those clusters. These environments are currently in IATT and working towards scale and ATO. Deliberately a single role: at current team size, build/release and run/operate are the same person's problem.
Release side: CDK stacks, Flux GitOps, Helm chart authoring, Envoy Gateway routes, EKS/IRSA, container builds, GitLab CI pipeline automation, and the local-to-AWS promotion path. Run side: the observability/tracing stack, cross-service tracing, runtime health, credential rotation, incident response, and the preflight/diagnose/QA loops.
This role directly supports OY2 Workstream 1 (Image Builder Pipeline — STIG-compliant AMI automation, Artifactory integration, GitLab CI deployment), Workstream 6 (Import Account Standardization — Baseline Account Pipeline, tagging compliance, centralized VPC endpoints), and the GenAI IDP deployment pipeline (deploying the IDP Solution to Customer non-production and production environments via IaC).
Basic Qualifications
- 3+ years of professional experience in SRE, DevOps, platform, or infrastructure engineering
- Experience operating Kubernetes (EKS) in production, including in DoD/IC classified environments
- Experience packaging and deploying applications with Helm (authoring and maintaining charts, not only consuming them)
- Experience with Flux (or an equivalent GitOps controller — e.g., Argo CD) driving continuous delivery of Helm releases
- Experience with AWS compute and managed services (e.g., EKS, RDS, S3, IAM/IRSA, EC2 Image Builder)
- Experience with infrastructure-as-code; TypeScript/CDK experience specifically, or demonstrated ability to work in a TypeScript IaC codebase
- Experience with production observability and distributed tracing (e.g., OpenTelemetry, Grafana/Tempo, or equivalent), used to diagnose failures from telemetry rather than guesswork
- Experience leading incident diagnosis and resolution, including identifying and confirming root cause before remediating
- Experience with GitLab CI/CD pipelines for build automation and deployment
- Proficiency in at least one scripting language (e.g., Python) for tooling and automation
Preferred Qualifications
- Experience building STIG-compliant AMI pipelines using EC2 Image Builder with DoD security baseline validation
- Experience with Artifactory integration for AMI/container image distribution across AWS Organizations
- Experience with cross-domain solutions (AWS Diode, CDS) and SIPRNet/JWICS environments
- Experience with AWS Organizations, SCPs, OU design, and multi-account governance
- Experience with certificate lifecycle management (ACM, Private CA) and automated renewal workflows
- Experience deploying or operating ML/LLM-serving infrastructure (e.g., Amazon Bedrock endpoints, SageMaker, model-inference endpoints) and reasoning about token/latency/cost from traces
- Experience with an ECR- or registry-based GitOps model (charts/images mirrored to a registry, controller reconciling from it)
- Experience diagnosing failed or stuck Helm/Flux reconciliations (HelmRelease not converging, drift between desired and live state)
- Experience with Envoy/gateway and certificate/TLS management in classified environments
- Proficiency using AI coding assistants as a daily driver, with the judgment to validate their output before it reaches a cluster
- Bachelor's degree in Computer Science or a related field, or equivalent practical experience
Job Type: Full-time
Pay: From $120,000.00 per year
Benefits:
- 401(k)
- 401(k) matching
- Dental insurance
- Flexible schedule
- Health insurance
- Life insurance
- Paid time off
- Referral program
- Retirement plan
- Tuition reimbursement
- Vision insurance
License/Certification:
- Security + or other Cyber/Security Cert? (Preferred)
Security clearance:
Work Location: Remote