MerchantE SLA Draft
Important: I could not reliably derive a complete MerchantE SLA with firm contractual commitments from the Confluence pages that were available. The content below is a best-effort operational SLA summary assembled from monitoring, incident-management, and SLO documentation. Items marked as Needs confirmation should be validated before external use or inclusion in a customer-facing agreement.
MerchantE SLA Overview
MerchantE operates a mature production support and observability model across payment gateway, authorization, settlement, file transfer, database, and supporting platform services. Based on the available Confluence documentation, MerchantE maintains continuous monitoring, defined P1 incident handling, production alerting through multiple tools, and published service objectives for key payment services.
What this page represents
This summary consolidates information from internal pages covering SLI/SLO targets, alerting, monitoring plans, and P1 incident procedures. It is suitable as a working draft or internal comparison artifact, but not as a final legal SLA without business and legal review.
Incident Response & Mitigation
1. Monitoring
MerchantE monitors critical production services on a continuous basis using a combination of external uptime checks, infrastructure monitoring, application telemetry, centralized logging, and alerting workflows.
Primary monitoring capabilities identified in Confluence include:
AWS Synthetic Canaries: External uptime and endpoint reachability monitoring for key production and certification endpoints, including Hosted Payments, Payment Gateway, Trident, token endpoints, and Virtual Terminal.
Icinga2 / AWS CloudWatch / Prometheus: Internal service and infrastructure monitoring for application load balancers, ingress objects, and application health.
CloudWatch Logs / Grafana / Splunk / Dashboards / Alarms: Centralized logging, alarm generation, and operational dashboards for AWS-hosted workloads.
Splunk: Log ingestion, operational alerting, and alert correlation across application, file transfer, infrastructure, database, and security domains.
PagerDuty: Escalation for selected production alerts and major incidents.
Monitoring scope includes:
Authorization applications such as PG, Trident, Hosted Payments, Split Funding, VT, and related services
Batch applications such as TEM, BP, ACH services, Notification Service, Product ID, PCI Compliance, and reporting-related services
Database infrastructure and replication/resource health
Settlement and outbound file delivery monitoring
CDC/data consistency monitoring for synchronized data flows
Operational strength: The monitoring estate is broad and layered. MerchantE uses both synthetic availability checks and internal telemetry, which reduces the chance of relying on a single detection method.
2. Notification
MerchantE has a documented P1 incident process, but the available pages do not define a full external notification matrix by severity in the same way Omise or Cybersource do.
What is documented:
For a P1 incident, the on-call Incident Manager is responsible for creating an MPI ticket and triggering PagerDuty for the Business War Room within 20 minutes of the initial P1 notification.
The Incident Manager must update the MPI ticket hourly with progress.
The Incident Manager acts as liaison between technical and business war rooms until service is restored and post-fix actions are complete.
What is not clearly documented in the pages reviewed:
Guaranteed customer-facing email notification times by priority
Status page commitments
External notification rules for P2, P3, or P4 incidents
Customer communication channel commitments for non-P1 incidents