Description
The position exists to close a known capability gap in log correlation, SLO and error-budget reporting, and early-warning detection.
Success in this role directly protects the performance-service-level track record AFS markets to clients and reduces the time it takes to detect and resolve production issues across the client portfolio.
Scope of Responsibility and Authority
The Observability Engineer owns the observability capability end to end across AFSVision HPC environments, including tool evaluation, log correlation architecture, and SLO/SLI definition.
The role has authority over day-to-day monitoring and observability engineering decisions, and escalates architectural changes, tooling investments, and cross-team resourcing needs to the Managing Director, System Services Group.
This position partners closely with Application Services, Database Administration, Network Engineering, and Information Security, and must design and operate all observability tooling within AFS’s data classification policy and applicable FFIEC, SOC 2, and NIST requirements, since the underlying telemetry can include client production data.
As AFS’s Next Gen Datacenter project moves workloads toward public cloud infrastructure, this scope extends to CI/CD pipeline and Kubernetes observability alongside the existing on-premises HPC environment.
Core Responsibilities
Build log correlation capability – design and stand up correlated, searchable log correlation across Linux, middleware, and AFSVision application logs, replacing ad-hoc, siloed log search.
Deliver SLO/SLI reporting – convert raw monitoring data into SLO/SLI and error-budget reporting tied directly to client SLAs, giving Operations and clients defensible reliability metrics.
Publish client-facing SLA reports – produce accurate, recurring performance and SLA reports for clients, reflecting measured system metrics.
Implement early-warning indicators – for the conditions that most often drive incidents, including batch run-time drift, resource saturation, and connection-tier health across the Linux and middleware tiers supporting AFSVision.
Lead the observability roadmap – evaluate and deploy tracing and synthetic monitoring inside the HPC trust boundary under FFIEC, SOC 2, and NIST controls, and extend that roadmap to the public cloud CI/CD pipeline and Kubernetes environments introduced by AFS’s Next Gen Datacenter project.
Maintain the core monitoring stack – keep SolarWinds, ManageEngine, and Site24x7 patched, compliant, and configured with accurate per-environment monitors and SLA dashboards.
Reduce alert noise – measurably cut MTTD and MTTR, lowering after-hours escalations and incident cost.
Partner across teams – work with Application Services, Database Administration, Network Engineering, and Information Security to correlate signals across the full stack during incident investigation.
Report on reliability – build and maintain dashboards and reliability reporting reviewed on a regular cadence with Operations and System Services Group leadership.
Support incident response – participate in on-call rotation and lead root cause analysis for production incidents involving monitoring or observability gaps.
Document standards – maintain observability standards, runbooks, and tool configuration documentation for adoption across the System Services Group, including monitoring practices for the Next Gen Datacenter’s cloud and Kubernetes environments as they come online.
Evaluate new tooling – pilot new observability and monitoring tools and present cost/benefit recommendations to leadership.
Maintain compliance – ensure all observability practices and tooling comply with AFS data classification policy and FFIEC, SOC 2, and NIST requirements.
First 12 Months Outcomes: What Success Looks Like
Success in the first year is defined by a functioning log correlation capability, SLO-based reliability reporting tied to client SLAs, and a measurable reduction in incident detection and resolution time, while meeting AFS’s compliance and quality standards.
Deliver log correlation – correlated, searchable log correlation live in production across AFSVision’s Linux, middleware, and application logs.
Establish SLO/SLI reporting – reviewed monthly with Operations and System Services Group leadership, tied directly to client SLAs.
Stand up client-facing SLA reporting – recurring, accurate SLA and performance reports delivered to clients on an agreed cadence.
Stand up early-warning indicators – for batch run-time drift, resource saturation, and connection-tier health, validated against at least one real incident pattern.
Reduce MTTD/MTTR – for production incidents by an agreed-upon target measured against the prior 12 months’ baseline.
Cut after-hours escalations – by reducing alert noise and improving signal quality.
Publish an observability roadmap – reviewed quarterly with System Services Group leadership, covering tracing, synthetic monitoring, and the observability requirements of the Next Gen Datacenter’s planned public cloud CI/CD pipeline and Kubernetes environments.
Pass compliance review – complete a full FFIEC/SOC 2/NIST audit cycle of observability tooling and data handling with zero unresolved critical findings.
Ways of Working
Proactive by default – looks for the next failure mode instead of waiting for the next page.
Data-driven communicator – translates raw telemetry into SLA/SLO language that Operations and clients can act on.
Security and compliance minded – treats client production telemetry as regulated data and builds controls in rather than adding them later.
Comfortable owning a 24x7x365 platform – including rotational on-call responsibility.
Collaborative – works across Operations, DBA, Application Engineering, Network Engineering, and Security rather than in isolation.
Profile
The right candidate is a hands-on builder who is more energized by standing up new observability capability than maintaining existing dashboards.
They think in terms of detection and prevention rather than reaction, and are comfortable translating infrastructure signals into business-relevant reliability and SLA metrics.
They operate independently but know when to escalate architectural or resourcing decisions, and they bring the discipline needed to work inside a regulated, audited environment.
Qualifications
We are looking for a hands-on observability engineer who can move AFS from reactive monitoring to proactive, correlated observability.
Demonstrated experience standing up log correlation and SLO-based reliability practices matters more than years in a title.
Experience designing and operating log correlation platforms at production scale, ideally across Linux, middleware, and application-tier logs.
Background maintaining enterprise monitoring tools (SolarWinds, ManageEngine, Site24x7, or similar), including patching, compliance, and dashboard configuration.
Demonstrated ability to build SLO/SLI frameworks and error-budget reporting tied to contractual SLAs.
Experience producing client-facing performance and SLA reporting.
Strong Linux systems background, with experience monitoring middleware and application logs for correlation and root cause analysis.
Cloud infrastructure experience (Azure preferred) and strong log analysis and correlation skills.
Comfort operating in a regulated, audited environment; able to work within FFIEC, SOC 2, and NIST constraints.
Track record of reducing MTTD/MTTR through improved detection, alerting, and incident response practices.
Exposure to public cloud CI/CD pipelines and Kubernetes monitoring is a plus, as AFS’s Next Gen Datacenter project migrates workloads off-premises.
Bachelor’s degree in Computer Science, Information Systems, or a related field, or equivalent hands-on experience.
Work Environment & Physical Demands
Work is performed primarily in an office setting during normal business hours, with periodic off-hours support required as part of an on-call rotation supporting a 24x7x365 production platform.
Flexible schedule options are available in accordance with AFS company policy.
This role requires high-frequency use of computer keyboards and monitors.
Position Description Disclaimer
To perform this job successfully, an individual must be able to perform each essential duty listed above satisfactorily.
Reasonable accommodation may be made to enable qualified individuals with disabilities to perform the essential functions of this role.
This position description is intended to provide guidelines for job expectations and is not an exhaustive list of all functions, responsibilities, skills, and abilities.
Additional functions and requirements may be assigned by supervisors as deemed appropriate.
This document does not constitute a contract of employment.
AFS reserves the right to modify this position description and assign tasks as business needs require.
Automated Financial Systems, LLC is an Equal Opportunity Employer.
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, veteran status, or any other characteristic protected by law.