We are looking for an experienced hands-on DevOps/SRE engineer to work with high-throughput bare-metal infrastructure and business-critical production systems.
REQUIREMENTS:
- 5+ years of hands-on experience in DevOps/SRE, including operating bare-metal infrastructure and high-throughput production systems
- Deep understanding of Linux internals, including CPU, memory management, file descriptors, system limits, IRQ affinity, and NUMA
- Strong knowledge of TCP/IP and the Linux networking stack, including the connection lifecycle, listen and accept queues, SYN backlog, TIME_WAIT, conntrack, ephemeral ports, retransmissions, and socket limits
- Experience diagnosing network saturation, packet loss, queue overflows, connection exhaustion, SYN backlog overflows, and abnormal traffic patterns
- Proficiency with production diagnostic tools such as
ss, tcpdump, perf, strace, and /proc
- Experience operating and tuning load balancers in high-throughput production environments
- Strong knowledge of Ansible, GitLab CI, Bash, and Git
- Experience with Prometheus or VictoriaMetrics, Grafana, and PromQL
- Experience participating in on-call rotations and resolving critical production incidents
- Ability to make technical decisions independently, drive complex initiatives, and share expertise with other engineers
- Ability to read and understand technical documentation in English
Deep Expertise in at Least One Area:
- Go in production: pprof, goroutine dumps, execution traces, and runtime metrics; diagnosing excessive allocations, GC pressure, goroutine leaks, lock contention, blocked I/O, and incorrect network timeouts; understanding GOGC, GOMEMLIMIT, GOMAXPROCS, connection pooling, and graceful shutdown. Ability to read Go code and provide developers with evidence-based hypotheses about performance and reliability issues
- Load balancing at scale: experience with systems processing more than 1 million requests per second
- Safe deployments across large server fleets: rolling and canary deployments, health checks, versioning, and rollback strategies
- Observability and performance: capacity planning, load testing, and systematic bottleneck analysis
WOULD BE A PLUS:
- Experience with latency-sensitive or real-time data processing systems
- Experience operating ClickHouse, Kafka, Redis, or Aerospike
- Experience with AdTech/OpenRTB
- Experience with MAAS/PXE, IPMI/BMC, or BGP
- Experience with Proxmox, FreeIPA, or nftables
- Practical experience using AI assistants in engineering workflows
- Spoken English at B1 level or higher
WE OFFER:
- Competitive salary
- A friendly and professional team
- Opportunities for professional growth and career development
- Paid vacation and sick leave
- Corporate English classes
- Medical insurance