Founded in 2007, Payliance is a trusted leader in payment processing — processing more than $63 billion annually, supporting 40,000+ merchant locations, and serving over 350 lending clients. We offer an all-in-one platform for real-time funding, payment processing, account verification, and recovery services, giving lenders the technology to operate efficiently and confidently.
What sets Payliance apart is our blend of modern technology, deep industry expertise, and a highly collaborative, people-first culture. Backed by Serent Capital, we're expanding our capabilities and delivering measurable results for clients across lending, e-commerce, collections, and gaming.
The Site Reliability Engineer (SRE) bridges software engineering and infrastructure operations, owning the reliability, scalability, and performance of Payliance's payment processing platform. You'll apply engineering discipline to operational problems — reducing toil, automating repetitive work, and building systems that are resilient by design.
This role is ideal for an engineer who can read and debug production .NET code, has hands-on AWS compute and RDS SQL Server depth, and brings the observability mindset needed to keep a high-transaction payment platform running at scale.
· Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues.
· Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents.
· Partner with application engineers to embed reliability into new feature design and deployment practices.
· Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early.
· Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit.
· Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments.
· Automate operational tasks, deployment pipelines, and disaster recovery procedures.
· Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work.
· Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup.
· Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly.
· Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains.
· Capacity plan and scale database infrastructure to support transaction volume growth.
· Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing.
· Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements.
· Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly.
· Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence.
· Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs.
· Configure and manage ALB/NLB routing, Route 53 DNS, and TLS certificate lifecycle via ACM.
· Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning.
· Support compliance and security controls relevant to a PCI-regulated payments environment.
· Participate in on-call rotation to respond to production incidents and drive swift resolution.
· Define and track error budgets; use them to balance velocity and reliability investment.
· Communicate status updates clearly during incidents and coordinate cross-functional response.
· Maintain and improve runbooks, escalation paths, and on-call health over time.
· Collaborate with platform engineering teams on architecture decisions and scalability requirements.
· Share observability and reliability best practices with application teams.
· Mentor engineers on SRE principles and operational excellence.
· Base Salary: $145,000 - $160,000
· Performance-based annual bonus.
· Medical, Dental, and Vision insurance.
· 401(k) with company match.
· Generous PTO plus paid company holidays.
· Company-paid life and long-term disability insurance.
· Paid parental leave.
Remote-first with collaboration across U.S. time zones. On-call duties are part of this role with rotation schedules designed to be sustainable and fairly distributed.
Payliance is an equal opportunity employer. We value diversity and strive to create an inclusive workplace for everyone. Discrimination or harassment of any kind — based on race, color, sex, religion, sexual orientation, gender identity, national origin, age, disability, genetic information, or pregnancy — is not tolerated. Reasonable accommodations are available throughout the application and employment process.
· 4+ years in SRE, DevOps, platform engineering, or a systems-focused software engineering role.
· C#/.NET engineering ability — can read, debug, and contribute to production code; experience diagnosing memory leaks, thread exhaustion, and GC pressure.
· AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each.
· RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution.
· Native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK.
· AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM.
· SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager.
· Strong scripting ability (PowerShell, Python, or Bash) for automation and operational tooling.
· Excellent communication skills and a collaborative, blameless engineering mindset.
· Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.
· Experience in fintech, payments, or high-transaction-volume regulated environments (PCI-DSS, SOC 2).
· Knowledge of ACH, card processing, or payment settlement workflows.
· Familiarity with chaos engineering or resilience testing (e.g., AWS Fault Injection Simulator).
· Experience with secrets management (AWS Secrets Manager, Parameter Store) and security scanning in CI/CD.
· Exposure to GitOps workflows and modern CI/CD practices.
· Experience working in a private equity-owned, venture-backed, or high-growth startup environment.
· Demonstrated ability to leverage AI tools to accelerate diagnostics, automate runbook creation, or improve observability workflows.
· Bachelor's degree in Computer Science, Engineering, or related field (or equivalent hands-on experience).
· Platform Uptime & SLO Attainment: Meeting or exceeding agreed service level objectives.
· MTTD & MTTR: Mean time to detection and resolution for production incidents.
· Toil Reduction: Measurable decrease in manual operational tasks through automation.
· Application Reliability: Reduction in .NET runtime issues (memory leaks, thread exhaustion) reaching production.
· Deployment Frequency & Stability: Supporting rapid iteration with minimal risk.
· Observability Coverage: Expansion of CloudWatch/X-Ray coverage and depth of incident insights.
· On-Call Health: Sustainable rotation with clear runbooks and blameless post-mortems.
· Team Impact: Mentorship and knowledge sharing that raises operational maturity across engineering.
Payliance
Payment processing technologies (ACH Processing, eCheck, RCC, Credit Card, Payment Gateway and Payment Recovery) for faster and more reliable payments.
Read more about the companyCopyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.