
- We are looking for a Senior DevOps / Platform Engineer to develop and operate GoMining's cloud platform.
- The primary focus of this role is GCP and Kubernetes, infrastructure and application delivery automation, and the reliability and security of production services.
- You will maintain and improve our existing infrastructure, contribute to the development of the next generation of our platform, and help product teams ship changes safely and efficiently.
Responsibilities
- Develop and operate infrastructure in GCP, including GKE, virtual machines, networking, load balancers, storage, and managed services.
- Manage Infrastructure as Code with Terraform: develop modules, manage state and dependencies, ensure reproducible changes, and identify configuration drift.
- Operate Kubernetes: upgrade clusters and platform components, configure autoscaling, workload placement, network policies, and resource allocation.
- Develop CI/CD and GitOps using GitLab CI and Flux / Argo CD, automating builds, checks, deployments, controlled releases, and rollbacks.
- Build observability across metrics, logs, distributed tracing, and actionable alerting. Define and improve SLIs/SLOstogether with engineering teams.
- Troubleshoot production incidents and performance issues, conduct root cause analysis, and eliminate recurring failure patterns.
- Implement least-privilege access, environment and service isolation, secure secrets management and rotation, and infrastructure audit controls.
- Maintain backup processes and regularly validate recovery of critical data and infrastructure according to agreed RPO/RTO.
- Help engineering teams prepare services for production, including health checks, resource configuration, telemetry, and deployment procedures.
- Optimize cloud infrastructure usage and costs while maintaining reliability requirements.
- Maintain infrastructure documentation, operational runbooks, and recovery procedures.
Requirements
- Strong hands-on experience independently operating production infrastructure and Kubernetes, including upgrades, troubleshooting, and disaster recovery.
- Strong experience with GCP, including GKE, Compute Engine, VPC, IAM, and Cloud Storage; good understanding of Cloud SQL, Cloud Logging, and Cloud Monitoring.
- Hands-on experience with Terraform, including modules, remote state, locking, importing existing resources, and safely applying infrastructure changes through CI.
- Deep knowledge of Kubernetes and containerization: Deployments, StatefulSets, Services, Ingress, storage, requests/limits, probes, autoscaling, RBAC, and NetworkPolicy.
- Experience with Helm and Kustomize.
- Strong Linux administration skills, preferably Debian/Ubuntu: processes, systemd, filesystems, disks, and resource troubleshooting.
- Strong networking fundamentals: TCP/IP, DNS, HTTP(S), TLS, routing, NAT, load balancing, and firewalls.
- Ability to troubleshoot connectivity issues across applications, Kubernetes clusters, networks, and cloud services.
- Experience building CI/CD pipelines, preferably with GitLab CI, and strong understanding of Git and GitOps principles.
- Understanding of runner permissions, secrets protection, and production deployment controls.
- Automation skills with Bash and Python, including maintainable infrastructure scripts and API integrations.
- Experience with monitoring, logging, and alerting using tools such as Prometheus, Grafana, or Cloud Monitoring.
- Practical experience operating PostgreSQL and understanding Redis and RabbitMQ, including availability, connections, replication or clustering, backup, and recovery.
- Strong understanding of cloud security principles: least privilege, service accounts, Workload Identity, secrets management, network segmentation, and access auditing.
- Ability to independently deliver infrastructure changes to a validated production result, explain technical trade-offs, and collaborate effectively with Engineering and Security teams.
Nice to Have
- Experience with Flux CD, Argo CD, Ansible, and HashiCorp Vault.
- Experience with OpenTelemetry and distributed tracing.
- Experience migrating infrastructure between clouds or Kubernetes clusters, including database migrations with limited downtime.
- Experience designing disaster recovery strategies and conducting recovery exercises.
- Cloud cost optimization and capacity planning experience.
- Experience with DigitalOcean, Hetzner, or AWS.
- Ability to read and troubleshoot JavaScript / TypeScript applications.
- Experience operating financial or payment services with strict requirements around access control, auditing, reliability, and data integrity.
Benefits
- Professional growth: support for courses, conferences, and English learning (up to 100% coverage).
- Work-life fit: remote or hybrid format with flexible hours across international teams.
- Paid leave: up to 20 vacation days + 8 company holidays + 5 personal days per year
- Recognition programs: structured performance reviews and team awards.
- Team culture: retreats in international locations (for example, company apartments in Cyprus).


