Senior Site Reliability Engineer - Observability
Location: Austin, TX Area (Remote-First)
Requirement: Candidates must be within commuting distance of Austin.
About the Role
A leading asset management firm is seeking a Senior Site Reliability Engineer to own and enhance its observability platforms. This role combines operational support with platform engineering, focused on improving reliability, scalability, and visibility across the technology organization.
Responsibilities
* Own the availability, performance, and support of observability platforms.
* Serve as an escalation point for monitoring and logging-related issues.
* Manage platform upgrades, capacity planning, performance tuning, and operational health.
* Partner with engineering teams to improve dashboards, alerting, and instrumentation.
* Build automation and self-service capabilities that reduce operational toil.
* Develop and maintain infrastructure-as-code and platform standards.
* Drive observability best practices and platform modernization initiatives.
* Maintain documentation, runbooks, and incident response processes.
Requirements
* 5+ years of experience in SRE, DevOps, Platform Engineering, or related roles.
* Strong hands-on experience with enterprise observability, logging, and monitoring platforms.
* Experience operating large-scale on-premises infrastructure and distributed systems.
* Strong Linux administration and troubleshooting skills.
* Proficiency in Python and scripting for automation.
* Experience with infrastructure automation and configuration management tools.
* Proven incident management and problem-solving capabilities.
* Strong communication skills and a passion for operational excellence.
Preferred Experience
* Modern observability and APM platforms.
* Distributed tracing and telemetry frameworks.
* Log collection and data pipeline technologies.
* Infrastructure-as-code and automation tooling.
* Financial services or other highly regulated environments.
