← All jobs
PN

Site Reliability Engineer (SRE) Manager

Plume NetworkLjubljana, Slovenia
Type
Full-time
Work setup
On-site
Experience
Mid
Posted
11 days ago

What engineering roles in crypto pay

205 salaries · our own data
median $202k$144k$255k

Most engineering roles in crypto pay between $144k and $255k, with a median of $202k.

As a Site Reliability Engineer Manager at Plume, you'll lead the NOC L2 team responsible for triaging, investigating, and resolving escalated network and platform incidents. You combine people leadership, operational ownership, and technical depth to build the systems, runbooks, and culture that make incident response faster, calmer, and more effective across the organization. Plume delivers services to over 60 million locations globally and has managed over 3 billion devices on its open, hardware-independent service delivery platform.

What you'll do

  • Lead, mentor, and grow a team of NOC L2 engineers responsible for escalated incident triage and resolution
  • Own on-call rotation structure, escalation policies, and incident response processes for the L2 team
  • Drive root-cause analysis and post-incident reviews, ensuring learnings translate into concrete process or system improvements
  • Partner with Engineering, Infrastructure, and Product teams to reduce recurring incident classes and improve platform reliability
  • Establish and refine runbooks, playbooks, and operational documentation to speed up incident resolution and reduce reliance on tribal knowledge
  • Monitor and report on key operational metrics (MTTD, MTTR, incident volume, escalation rates) to leadership
  • Manage team schedules, coverage, and on-call rotations to ensure 24/7 operational readiness
  • Hire, coach, and develop engineers on the team through regular 1:1s, performance reviews, and career development planning
  • Act as an escalation point for the most critical or ambiguous incidents, providing hands-on technical guidance when needed
  • Collaborate with L1 NOC leadership to ensure smooth escalation handoffs and continuous improvement of triage criteria
  • Drive a culture of blameless postmortems and continuous operational learning

What you bring

  • 4+ years of experience in network operations, infrastructure, or site reliability roles, with at least 2+ years in a people management or team lead capacity
  • Strong understanding of networking fundamentals including TCP/IP, DNS, routing, firewalls, and VPNs, plus troubleshooting methodology
  • Proven experience owning incident response processes, including on-call rotation design and escalation management
  • Experience with monitoring, alerting, and observability tools such as Grafana, Prometheus, Datadog, PagerDuty, and Splunk
  • A strong track record of driving operational improvements that measurably reduce incident volume or resolution time
  • Excellent communication skills, able to translate technical incidents into clear updates for both technical and non-technical stakeholders
  • Comfort operating in a 24/7 operational environment, including managing coverage across shifts and time zones

Nice to have

  • Experience in telecom, ISP, networking hardware, or connected-device industries
  • Familiarity with cloud infrastructure such as AWS, GCP, or Azure, and container orchestration including Kubernetes and Docker
  • Experience with automation and scripting to reduce manual operational toil using Python, Bash, or similar languages
  • Prior experience scaling a NOC or SRE team through significant growth
  • ITIL or similar operational framework certification or experience

About Plume

Plume is the creator of the only open, hardware-independent, cloud-controlled experience platform for ISPs and their subscribers, partnering with over 400 ISP customers including Charter, Liberty Global, and J:COM. The company's OpenSync framework enables ISPs to decouple service offerings from hardware and rapidly deliver new services over a multi-vendor, open-platform architecture.

Site Reliability Engineer (SRE) Manager | CryptoJobsHQ