Log inPost your job
← All jobs
GA

Vice President Site Reliability Engineering (Data Centers)

GalaxyRemote
EngineeringLead / HeadHybrid

As Vice President of Site Reliability Engineering (Data Centers) at Galaxy, you lead a specialized SRE team overseeing automation toolsets and the systems they interact with. You combine strategic vision with hands-on technical expertise in infrastructure automation at scale, treating infrastructure as a product. You have a proven track record managing complex hybrid environments and building self-service platforms that enhance engineering velocity and system stability across Galaxy's global data center footprint.

What you'll do

  • Oversee a specialized SRE team focused on the design, deployment, and maintenance of automation toolsets and the systems they interact with
  • Establish and enforce standards for Infrastructure as Code to ensure consistent, repeatable, and secure deployments across the entire infrastructure ecosystem, with strong proficiency in Terraform
  • Lead the strategy for automated configuration and state management, ensuring Ansible playbooks and Packer image pipelines are optimized for Windows, Linux, and ESXi platforms
  • Manage monitoring and health of the automation platforms themselves, implementing SLIs and SLOs to ensure high availability and performance of the tools that build the servers
  • Drive the automated lifecycle of both physical and virtual assets, from initial template creation and deployment through automated patching, scaling, and decommissioning
  • Lead the development of custom scripts and internal providers in Python, Go, PowerShell, and Bash to provide better insights and tooling for your systems
  • Collaborate with the broader datacenter team to foster workflows and facilitate needs across the organization
  • Analyze system behavior and resource utilization in virtual environments to optimize the performance of automated deployments
  • Provide technical guidance and career mentorship to SREs, fostering a culture of automate-first and continuous improvement

What you bring

  • 6-10 years of experience in infrastructure, SRE, or DevOps, specifically focused on infrastructure automation at scale
  • Deep proficiency with Terraform, including providers, modules, and state management
  • Deep proficiency with Ansible, including roles, playbooks, and Tower/AWX
  • Hands-on experience with image creation tools such as Packer and Ansible to build standardized, hardened images for both Windows and Linux in hybrid environments
  • Strong experience managing and automating virtual platforms such as VMware (vSphere/vCenter) as well as cloud providers such as Azure and AWS
  • High-level scripting skills in Python, Go, PowerShell, and Bash
  • Experience with observability tools such as Splunk, ELK, Prometheus, or Grafana to monitor infrastructure health and automation telemetry
  • A good understanding of network topology and design, with experience on platforms such as Juniper Networks or Palo Alto
  • Strong mastery of Git, including branching strategies and PR workflows, and CI/CD platforms such as Jenkins, GitLab CI, or GitHub Actions
  • Equal comfort managing, troubleshooting, and tuning performance for both Windows Server and Linux

Nice to have

  • Previous work experience that includes notable periods of team leadership or management
  • Experience with IAM platforms such as Entra ID, Active Directory, and Okta
  • Experience with storage solutions, both block-based and object-based, hosted either on-premises (HP Alletra, EMC, DDN) or in the cloud (S3, Azure Blob)
  • Storage backup and disaster recovery administration and management with Commvault or Veeam

About Galaxy

Galaxy is a global leader in digital assets and data center infrastructure. The company operates an institutional digital assets platform spanning trading, investment banking, asset management, staking, self-custody, and tokenization technology, and invests in and operates cutting-edge data center infrastructure to power AI and high-performance computing. Headquartered in New York City, Galaxy has offices across North America, Europe, the Middle East, and Asia.

What engineering roles in crypto pay

383 salaries · our own data
median $205k$157k$250k

Most engineering roles in crypto pay between $157k and $250k, with a median of $205k.

Vice President Site Reliability Engineering (Data Centers) | CryptoJobsHQ