Logo-of-Payliance-hiring-for-jobs-in-US-on-GrabJobs

Site Reliability Engineer (SRE)

salary Salary :

$145,000 - 160,000 yearly

icon building Company : Payliance
icon briefcase Job Type : Full Time
icon remote-alt Remote / Work from Home

Job Description - Site Reliability Engineer (SRE)


  

About Payliance

Founded in 2007, Payliance is a trusted leader in payment processing — processing more than $63 billion annually, supporting 40,000+ merchant locations, and serving over 350 lending clients. We offer an all-in-one platform for real-time funding, payment processing, account verification, and recovery services, giving lenders the technology to operate efficiently and confidently.

What sets Payliance apart is our blend of modern technology, deep industry expertise, and a highly collaborative, people-first culture. Backed by Serent Capital, we're expanding our capabilities and delivering measurable results for clients across lending, e-commerce, collections, and gaming.

About the Role

The Site Reliability Engineer (SRE) bridges software engineering and infrastructure operations, owning the reliability, scalability, and performance of Payliance's payment processing platform. You'll apply engineering discipline to operational problems — reducing toil, automating repetitive work, and building systems that are resilient by design.

This role is ideal for an engineer who can read and debug production .NET code, has hands-on AWS compute and RDS SQL Server depth, and brings the observability mindset needed to keep a high-transaction payment platform running at scale.

What You'll Do

.NET Application Reliability

· Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues.

· Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents.

· Partner with application engineers to embed reliability into new feature design and deployment practices.

· Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early.

AWS Compute & Infrastructure

· Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit.

· Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments.

· Automate operational tasks, deployment pipelines, and disaster recovery procedures.

· Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work.

RDS SQL Server Operations

· Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup.

· Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly.

· Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains.

· Capacity plan and scale database infrastructure to support transaction volume growth.

Observability & Monitoring

· Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing.

· Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements.

· Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly.

· Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence.

Networking & Security

· Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs.

· Configure and manage ALB/NLB routing, Route 53 DNS, and TLS certificate lifecycle via ACM.

· Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning.

· Support compliance and security controls relevant to a PCI-regulated payments environment.

Incident Response & On-Call

· Participate in on-call rotation to respond to production incidents and drive swift resolution.

· Define and track error budgets; use them to balance velocity and reliability investment.

· Communicate status updates clearly during incidents and coordinate cross-functional response.

· Maintain and improve runbooks, escalation paths, and on-call health over time.

Cross-Functional Partnership

· Collaborate with platform engineering teams on architecture decisions and scalability requirements.

· Share observability and reliability best practices with application teams.

· Mentor engineers on SRE principles and operational excellence.

  

Compensation & Benefits

· Base Salary: $145,000 - $160,000

· Performance-based annual bonus.

· Medical, Dental, and Vision insurance.

· 401(k) with company match.

· Generous PTO plus paid company holidays.

· Company-paid life and long-term disability insurance.

· Paid parental leave.

Work Environment

Remote-first with collaboration across U.S. time zones. On-call duties are part of this role with rotation schedules designed to be sustainable and fairly distributed.

Equal Employment Opportunity

Payliance is an equal opportunity employer. We value diversity and strive to create an inclusive workplace for everyone. Discrimination or harassment of any kind — based on race, color, sex, religion, sexual orientation, gender identity, national origin, age, disability, genetic information, or pregnancy — is not tolerated. Reasonable accommodations are available throughout the application and employment process.


Requirements

  

What You'll Bring

Required Qualifications

· 4+ years in SRE, DevOps, platform engineering, or a systems-focused software engineering role.

· C#/.NET engineering ability — can read, debug, and contribute to production code; experience diagnosing memory leaks, thread exhaustion, and GC pressure.

· AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each.

· RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution.

· Native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK.

· AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM.

· SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager.

· Strong scripting ability (PowerShell, Python, or Bash) for automation and operational tooling.

· Excellent communication skills and a collaborative, blameless engineering mindset.

· Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.

Preferred Qualifications

· Experience in fintech, payments, or high-transaction-volume regulated environments (PCI-DSS, SOC 2).

· Knowledge of ACH, card processing, or payment settlement workflows.

· Familiarity with chaos engineering or resilience testing (e.g., AWS Fault Injection Simulator).

· Experience with secrets management (AWS Secrets Manager, Parameter Store) and security scanning in CI/CD.

· Exposure to GitOps workflows and modern CI/CD practices.

· Experience working in a private equity-owned, venture-backed, or high-growth startup environment.

· Demonstrated ability to leverage AI tools to accelerate diagnostics, automate runbook creation, or improve observability workflows.

· Bachelor's degree in Computer Science, Engineering, or related field (or equivalent hands-on experience).

How Your Success Will Be Measured

· Platform Uptime & SLO Attainment: Meeting or exceeding agreed service level objectives.

· MTTD & MTTR: Mean time to detection and resolution for production incidents.

· Toil Reduction: Measurable decrease in manual operational tasks through automation.

· Application Reliability: Reduction in .NET runtime issues (memory leaks, thread exhaustion) reaching production.

· Deployment Frequency & Stability: Supporting rapid iteration with minimal risk.

· Observability Coverage: Expansion of CloudWatch/X-Ray coverage and depth of incident insights.

· On-Call Health: Sustainable rotation with clear runbooks and blameless post-mortems.

· Team Impact: Mentorship and knowledge sharing that raises operational maturity across engineering.


Original job Site Reliability Engineer (SRE) posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

About the Company

Payliance

Payment processing technologies (ACH Processing, eCheck, RCC, Credit Card, Payment Gateway and Payment Recovery) for faster and more reliable payments.

Read more about the company

Similar Site Reliability Engineer Jobs in the US

GrabJobs is the no1 job portal in the US, connecting you to thousands of jobs fast! Find the best jobs in the US, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.