Victoria Metrics:
Technical Skills
Hands-on operation of Victoria Metrics in production — VM Cluster topology (VM Insert, VM Storage, VM Select), vm agent scrape and stream aggregation, VM Auth multi-tenant routing, VM Alert and VM Alert manager rule design, and Metrics QL query authoring. Ability to govern active time series cardinality, configure retention and down sampling, and design multi-cluster federation with cross-cluster deduplication and write-path failover.
Experience Profile:
Verifiable production VM Cluster deployments handling sustained, high-cardinality workloads — not evaluation or sandbox only. Evidence of diagnosing and remediating a cardinality problem in a live environment. Operational experience of VM Auth-based per-tenant data isolation beyond authentication alone. Familiarity with the Victoria Metrics Operator on Kubernetes and its upgrade lifecycle. Intermediate accepted where Expert-level Observability is demonstrated.
Architecture Leadership:
Ability to serve as the single accountable design authority for a complex technical platform — producing High-Level Designs, Low-Level Designs, Architecture Decision Records, sequence diagrams, and interface contracts that engineering teams build from directly. Comfortable presenting architecture decisions and their trade-offs to mixed audiences from senior engineers to executive sponsors, adjusting depth without losing accuracy.
Has held a Principal or Staff Architect role where design and build were explicitly separated. Maintained an ADR record throughout a programme delivery. Can name a design decision they were challenged on, the trade-offs documented, and how the decision held up through delivery. Evidence of architecture governance participation: design review boards, security assurance gates, and formal approval before engineering investment began.
Kubernetes & OpenShift
Kubernetes administration and troubleshooting across StatefulSets, PersistentVolumeClaims, StorageClasses, Operators, RBAC, and NetworkPolicy. Red Hat OpenShift-specific competency: Security Context Constraints, OCP upgrade path management, OLM operator lifecycle, and the material differences from cloud-managed Kubernetes that affect stateful, high-throughput workloads. Multi-cluster topology design including hub and spoke architectures and cross-cluster connectivity.
Operational experience on Red Hat OpenShift in an on-premises or bare-metal enterprise environment — not exclusively cloud-managed Kubernetes. Has configured SCCs for a stateful workload without cluster-admin workarounds. Managed an OCP cluster through at least one version upgrade including Operator compatibility validation. Experience with OpenShift ACM or equivalent for policy and workload deployment across a cluster fleet.
GitOps & CI/CD
GitOps-native delivery model using ArgoCD or FluxCD as the reconciliation engine — all cluster state managed in Git, no manual production changes permitted. Kustomize overlay strategy for multi-environment and multi-tenant deployments: base resource definitions with environment-specific patches. GitLab CI/CD pipeline design including Kubernetes manifest validation gates, custom resource health checks, and environment promotion from lab through staging to production.
Has operated a GitOps-exclusive delivery model in production where ArgoCD or Flux managed all cluster state and direct kubectl commands to production were prohibited. Designed a Kustomize overlay structure for a multi-cluster platform workload — not a single-cluster application. Built a GitLab CI pipeline with schema validation and promotion gates blocking invalid configurations before staging. Evidence of GitOps discipline applied to operator-managed custom resources, not only standard Deployments.
Pay: $60.00 - $65.00 per hour
Application Question(s):
Work Location: Remote
Read authentic reviews with a Glassdoor account. Only apply to jobs you love.