I
Infobip

Site Reliability Engineer

Middle Hybrid Zagreb, Croatia Full-time Posted 1 month ago

About the role

Working at Infobip means being part of something truly global. With 75+ offices across six continents, we’re not just building technology — we’re shaping how more than 80% of the world connects and communicates.
As employees, we take pride in contributing to the world’s largest and only full-stack cloud communication platform. But it’s not just what we do, it’s how we do it: with curiosity, passion, and a whole lot of collaboration.
We operate with an AI-first mindset, embedding intelligent tools into our daily workflows to work smarter and more efficiently. Every role here benefits from and contributes to this approach.
If you're looking for meaningful work and challenges that grow you in a culture where people show up with purpose, this is your opportunity.
Let’s build what’s next, together.
Why is this position important at Infobip?
We are looking for engineers who enjoy solving problems, care deeply about quality, and take a data-driven, analytical approach to their work. The role we are hiring for is Site Reliability Engineer (SRE), with a primary focus on driving reliability initiatives and promoting reliability practices across the organization.
For this position, we are specifically seeking an SRE with a strong development background, ideally in backend or full-stack engineering, who is comfortable with scripting and building internal tools.
An SRE in this role understands client use cases, usage patterns, and business impact. They help establish and improve incident management processes, lead post-incident reviews, provide incident metadata and trend analysis, and act as advisors on follow-up actions. Their focus is on promoting a systematic approach to problem-solving and driving long-term solutions.
The team for which we are hiring has a diverse skill set as they come from different backgrounds, from system and software engineering, product management and testing. Over the course of long and varied careers, they have built a broad and valuable skill set that makes them exceptionally strong at handling problems and incidents. Today, they operate as a center of excellence, helping Infobip solve a wide range of complex reliability challenges.
SREs at Infobip are:
1. Owners of incident management and its lifecycle
Aligning incident management processes with the rest of the company
Streamlining workflows and automating processes
Driving adoption and monitoring process implementation
Creating objective, actionable incident reports
2. Advisors and experts on reliability topics
Supporting onboarding and education on reliability practices
Driving a reliability-focused culture and mindset
Promoting a blameless incident culture
Identifying risks and encouraging a systematic approach to problem-solving and long-term improvements
3. Partners in defining, developing, monitoring, and maintaining SLOs and SLIs
Helping teams set meaningful reliability targets for their products
4. Providers of a client-centric perspective
Bringing the client point of view into incident handling and product reliability discussions
5. Contributors to faster incident response
Helping shorten response times across detection, engagement, and resolution through troubleshooting support and continuous improvements
6. Builders of internal tooling and automation
Reducing operational toil and accelerating incident response
7. Providers of objective quality insights
Raising awareness of improvement areas based on historical data
Delivering actionable insights through quality trends and metrics
Guiding teams when metrics fall outside expected baselines

Responsibilities

Discovering problems, defining tasks, and solving them with guidance from more senior engineers
Designing and implementing automation and tooling, such as scripts, services, and dashboards, to improve incident detection and response, reduce manual work, and provide better reliability insights
Monitoring incidents in a production environment and supporting others during complex incident response situations
Troubleshooting platform-wide issues by understanding the bigger picture and guiding teams toward resolution of reliability-related problems
Collaborating with product and development teams to integrate reliability improvements into code and architecture
Continuously learning about Infobip’s development processes, technologies, system architecture, platform, products, and client segmentation
Actively participating in incident reviews and learning from incidents
Communicating problems and solutions at the right level of detail for different audiences involved in incident management
Sharing problem-solving knowledge within the team and the broader domain

Apply for this role

Send a short intro — the hiring team will review and respond.

I
Product

Croatian-founded global cloud communications platform (CPaaS) providing messaging, voice, chat and customer-engagement APIs to enterprises, with major engineering hubs in Vodnjan and Zagreb.

Location
Zagreb, Croatia
Team size
3500+
Remote policy
Hybrid
View company profile

Overview

Seniority
Middle
Remote policy
Hybrid
Location
Zagreb, Croatia
Employment
Full-time
Posted
Jul 27, 2026

Similar jobs