Lead Engineer – Reliability & Observability

Bengaluru, Karnataka, India | Full-time

Apply

About the Role

We are looking for a Lead Engineer to own and significantly advance our Reliability & Observability platform. We already have an established observability stack, including ELK, Grafana, Loki and monitoring dashboards. As our multi-tenant platform scales across multiple deployment sites and environments, you will drive the next phase of its evolution, making it easier for engineering teams to understand, monitor and improve the reliability of the systems they build.

This is a hands-on engineering leadership role. You will own and evolve our observability stack, establish reliability engineering standards and drive automation across our production environment. You will work closely with product/application engineering, cloud infrastructure and security teams to build resilient, self-operating systems across an increasingly complex distributed infrastructure.

 

What You'll Own

  • Observability Platform: Own, scale and continuously improve our existing metrics, logging, distributed tracing, dashboards and alerting infrastructure to support increasing platform scale, multiple deployment sites and multi-tenant environments.
  • Reliability Engineering: Establish SLO frameworks, error-budget reporting, production-readiness standards and reliability engineering best practices.
  • Developer Enablement: Build self-service observability capabilities, standardized instrumentation, reusable dashboards and actionable alerting for engineering teams.
  • Incident Management: Establish incident-management tooling, escalation processes, runbooks and post-incident review practices. Drive improvements that prevent recurring incidents.
  • Automation: Eliminate operational toil through automation, intelligent alerting, automated remediation and resilient system design.
  • Technical Leadership: Own the architecture and roadmap for Reliability & Observability, mentor engineers and collaborate with other engineering teams on cross-cutting reliability challenges.

Application teams own their services, and cloud infrastructure engineers own the underlying infrastructure. Your role is to provide the shared capabilities, standards and engineering expertise that help them operate reliably.

 

What We're Looking For

  • 7+ years of software engineering, platform engineering or SRE experience, including operating business-critical production systems at high scale and complexity.
  • Strong experience with observability technologies such as Elasticsearch, Logstash, Kibana (ELK), Grafana, Loki, OpenTelemetry, Jaeger or equivalent platforms.
  • Hands-on experience with AWS, Kubernetes, Linux, networking and distributed systems.
  • Strong programming and scripting skills in Go, Java, Python or similar languages, with an automation-first mindset.
  • Experience designing, scaling and improving production observability platforms, effective monitoring, alerting, SLOs and incident-management practices.
  • A solid understanding of distributed-system failure modes, scalability, fault tolerance and production debugging.
  • Experience with infrastructure as code, CI/CD and cloud-native architectures.
  • Ability to drive technical initiatives independently, make sound architectural decisions and influence engineering teams without relying on organizational authority.

 

Good to Have

  • Experience operating high-availability, multi-tenant SaaS platforms, particularly in fintech or other regulated environments.
  • Experience designing observability architectures for geographically distributed infrastructure, multiple Kubernetes clusters and isolated tenant environments.
  • Experience with large-scale telemetry pipelines, observability cost optimization and high-cardinality metrics.
  • Experience with chaos engineering, resilience testing and automated incident remediation.
  • Experience building internal developer platforms or self-service engineering tools.

Above all, we are looking for an engineer who enjoys building reliable platforms, not just maintaining monitoring tools.

 

Why Join Us

  • Build a high-impact platform in the financial services domain
  • Opportunity to work with some of the best brains in fintech and a strong product engineering team
  • Grow exponentially by working in small and transparent teams and influence tech culture
  • Increase your geek quotient by attending meetups and conferences.

 

About Cybrilla

Cybrilla is a financial infrastructure company aiming to disrupt the way mutual funds work. We work with some of the largest financial institutions as well as some of the fastest growing fintechs in the country. We are decentralizing distribution and changing the face of the industry.  We are a SEBI registered RTA and are co-authoring the Mutual Fund protocol on ONDC. More details: https://cybrilla.com

Our product, Fintech Primitives (FP) is an API platform that provides solutions to the problem statements of the Indian Mutual Fund domain. The APIs abstract domain, regulatory, and technical complexities to enable customers to build different distribution use cases in a short time. More details: https://docs.fintechprimitives.com