Site Reliability Engineer at BCW Group
- 8+ years of experience in IT infrastructure, cloud engineering, DevOps, and Site Reliability Engineering - Developed strong expertise in building, automating, monitoring, and maintaining highly available infrastructure across AWS, GCP, Kubernetes, Linux, and bare-metal environments. - Work extensively with Web3 infrastructure, managing high-throughput RPC, Validator, and Archive nodes across networks including Ethereum, Cosmos, ZK, and Solana. - Core responsibilities include incident management, performance and sync optimization, validator safety, automated failover, infrastructure provisioning, and maintaining high service availability. - Hands-on experience with modern DevOps and GitOps technologies including Terraform, Ansible, Docker, Kubernetes, Helm, GitLab CI/CD, and ArgoCD. - Observability and reliability - DataDog, Prometheus, Mimir, Grafana, and AppDynamics, developing monitoring and alerting systems that support proactive incident detection and operational excellence