MCPNew: Mokaru MCP server is live
Unifonic

Unifonic

Senior Site Reliability Engineer

Company

Unifonic

Role

Senior Site Reliability Engineer

Job type

Full-time

Found on Mokaru

4 days ago

Share this job

Salary

Not disclosed by employer

Job description

Proudly voted a Great Place to Work®, we are a dynamic startup in the CPaaS (Communication Platform as a Service) space that is revolutionizing the way businesses communicate. Our team is made up of 500 energetic and passionate Unifones who are dedicated to delivering the best possible experience to 5000+ customer-centric companies. We pride ourselves on our fun and collaborative work environment, where creativity and new ideas are constantly encouraged. As shareholders in the business, we’re so much more than a group of passionate communicators. We are Unifones. Join our team and be a part of something big! Meet the team! Our Engineering team is responsible for designing, developing, and maintaining the systems and technologies that drive Unifonic’s solutions. We work closely with other departments to ensure our products and services meet the needs of our customers. If you are passionate about technology and are excited about working on cutting-edge communication and engagement solutions, we want you on our team. As a Senior Infrastructure Engineer in the Production Operations (Live) team you will be responsible for enhancing system reliability, scalability, and resilience. As part of our elite SRE team, you'll drive continuous improvement across our cloud infrastructure and ensure the consistent high performance of our distributed messaging platforms. Help us shape the future of communication by: Owning the reliability, uptime, and scalability of critical production services 24/7. Participating in the on-call rotation to respond to incidents, troubleshoot live production issues, and lead post-incident analysis 24/7. Building robust operational playbooks, escalation paths, and improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). Ensuring operational excellence by proactively detecting and addressing reliability risks through SLO monitoring, chaos testing, and capacity planning. Automating operational tasks to minimize human intervention. Being available at night during the usual non-working hours of the rest of the team according to the on-call schedule is a MUST. Architecting, implementing, and managing infrastructure across AWS, Oracle Cloud Infrastructure (OCI), and OpenStack environments. Optimizing cloud resources to balance performance, security, and cost-efficiency. Managing Kubernetes clusters (EKS, OKE, Rancher RKE2), ensuring scalability, availability, and robust performance. Deploying advanced containerization strategies and troubleshooting. Managing and optimizing high-performance messaging and caching systems including Kafka, RabbitMQ, and Redis. Ensuring efficient, reliable message and data delivery critical to Unifonic's SMS and distributed systems. Managing and optimizing production-grade MySQL and PostgreSQL databases. Ensuring high availability, performance tuning, backups, and recovery processes for critical databases. Leading the planning and execution of comprehensive disaster recovery strategies. Developing and maintaining robust business continuity plans. Implementing advanced observability solutions (Prometheus, Grafana, CloudWatch). Defining, measuring, and enforcing Service Level Objectives (SLOs) and Service Level Indicators (SLIs) in alignment with SRE best practices. Proactively identifying issues, minimizing downtime, and enhancing system transparency. Driving automation initiatives using Terraform, Helm, Jenkins, Tekton or GitLab CI/CD. Streamlining deployment pipelines and reduce manual intervention through innovative automation. Integrating security best practices into infrastructure and application layers. Performing regular audits ensuring compliance and robust security posture. Collaborating with cross-functional teams (engineering, product, QA) to foster SRE culture. Mentoring junior engineers, enhancing team capabilities and promoting knowledge sharing.

Resume ExampleCover Letter Example

Explore more