observability engineer
Apptoza Inc.
Role: Observability Engineer / Banking Domain Work Model: Hybrid – 4 days onsite per week Key Responsibilities
Design, implement, and maintain enterprise observability solutions covering
metrics, logs, traces, events, and application performance .
Develop and manage monitoring and alerting strategies for critical banking applications and infrastructure.
Implement
OpenTelemetry
and distributed tracing across microservices and cloud-native applications.
Work with platforms such as
Splunk, Dynatrace, Grafana, Prometheus, Elastic/ELK , or similar observability technologies.
Create and maintain dashboards, alerts, service-level indicators (SLIs), and service-level objectives (SLOs).
Monitor applications, APIs, databases, Kubernetes/container environments, and cloud infrastructure.
Integrate observability tooling into
CI/CD and DevOps pipelines .
Automate monitoring, alerting, and operational processes using
Python, PowerShell, Bash, or similar scripting languages .
Analyze production incidents and performance issues using telemetry data and provide actionable recommendations.
Collaborate with application developers, DevOps, SRE, infrastructure, security, and production support teams.
Establish observability standards, reusable patterns, and best practices across engineering teams.
Participate in incident response, root-cause analysis, and post-incident reviews.
Implement proactive monitoring to identify performance degradation and potential production issues before they impact customers.
Ensure monitoring and observability practices align with banking security, compliance, audit, and risk requirements.
Support capacity planning, performance optimization, and reliability engineering initiatives.
Required Skills & Experience
10+ years of experience in
Observability, SRE, DevOps, Production Engineering, or Monitoring Engineering .
Strong understanding of
observability concepts: logs, metrics, traces, events, alerting, SLIs/SLOs, and distributed tracing .
Hands-on experience with one or more major observability platforms:
Splunk
Dynatrace
Grafana
Elastic Stack / ELK
AppDynamics
New Relic
Strong experience with
OpenTelemetry
and/or modern distributed tracing.
Experience monitoring
microservices, REST APIs, Kubernetes, Docker, and cloud environments .
Strong knowledge of
Linux/Unix
environments.
Experience with scripting/programming using
Python, Bash, PowerShell, or Java/Groovy .
Experience with
AWS, Azure, or GCP
observability and monitoring services.
Understanding of
CI/CD, Git, Jenkins, GitHub Actions, GitLab, or Azure DevOps .
Knowledge of databases, messaging platforms, APIs, and distributed systems.
Strong troubleshooting and analytical skills.
Experience working in
production/24x7 enterprise environments .
Previous experience supporting
banking, financial services, payments, capital markets, or insurance
environments is highly preferred.
Understanding of production controls, change management, incident management, and operational risk.
Familiarity with environments requiring strong
security, auditability, resiliency, availability, and regulatory compliance .
Experience supporting high-volume, business-critical applications is an advantage.
#J-18808-Ljbffr
Design, implement, and maintain enterprise observability solutions covering
metrics, logs, traces, events, and application performance .
Develop and manage monitoring and alerting strategies for critical banking applications and infrastructure.
Implement
OpenTelemetry
and distributed tracing across microservices and cloud-native applications.
Work with platforms such as
Splunk, Dynatrace, Grafana, Prometheus, Elastic/ELK , or similar observability technologies.
Create and maintain dashboards, alerts, service-level indicators (SLIs), and service-level objectives (SLOs).
Monitor applications, APIs, databases, Kubernetes/container environments, and cloud infrastructure.
Integrate observability tooling into
CI/CD and DevOps pipelines .
Automate monitoring, alerting, and operational processes using
Python, PowerShell, Bash, or similar scripting languages .
Analyze production incidents and performance issues using telemetry data and provide actionable recommendations.
Collaborate with application developers, DevOps, SRE, infrastructure, security, and production support teams.
Establish observability standards, reusable patterns, and best practices across engineering teams.
Participate in incident response, root-cause analysis, and post-incident reviews.
Implement proactive monitoring to identify performance degradation and potential production issues before they impact customers.
Ensure monitoring and observability practices align with banking security, compliance, audit, and risk requirements.
Support capacity planning, performance optimization, and reliability engineering initiatives.
Required Skills & Experience
10+ years of experience in
Observability, SRE, DevOps, Production Engineering, or Monitoring Engineering .
Strong understanding of
observability concepts: logs, metrics, traces, events, alerting, SLIs/SLOs, and distributed tracing .
Hands-on experience with one or more major observability platforms:
Splunk
Dynatrace
Grafana
Elastic Stack / ELK
AppDynamics
New Relic
Strong experience with
OpenTelemetry
and/or modern distributed tracing.
Experience monitoring
microservices, REST APIs, Kubernetes, Docker, and cloud environments .
Strong knowledge of
Linux/Unix
environments.
Experience with scripting/programming using
Python, Bash, PowerShell, or Java/Groovy .
Experience with
AWS, Azure, or GCP
observability and monitoring services.
Understanding of
CI/CD, Git, Jenkins, GitHub Actions, GitLab, or Azure DevOps .
Knowledge of databases, messaging platforms, APIs, and distributed systems.
Strong troubleshooting and analytical skills.
Experience working in
production/24x7 enterprise environments .
Previous experience supporting
banking, financial services, payments, capital markets, or insurance
environments is highly preferred.
Understanding of production controls, change management, incident management, and operational risk.
Familiarity with environments requiring strong
security, auditability, resiliency, availability, and regulatory compliance .
Experience supporting high-volume, business-critical applications is an advantage.
#J-18808-Ljbffr
Reference: WJ-3875_13223601