M01 - Cloud Logging & Data Platform Engineer

FPT Asia Pacific Pte LtdSingapore, Singapore, Singapore
Activeverified 6h ago

Job description

Overview We are looking for a Logging & Data Platform Engineer to design, build, and operate the logging and operational data-platform capabilities supporting the future SSOE platform. You will build the platform that collects, transports, stores, indexes, searches, and serves operational data across MOE's technology environment, spanning on-premise infrastructure, networks, applications, GCC, AWS, Azure, and hybrid environments. What You Will Be Working On As a Logging & Data Platform Engineer, you will build the shared platform capabilities that enable engineering and operations teams to reliably collect and use operational data at scale.You will work across logging, telemetry ingestion, data movement, storage, search, retention, and platform integration.The role requires an engineer who understands both traditional enterprise infrastructure and modern cloud-native architectures and can design solutions that work securely and reliably across environment boundaries.You will work closely with the Observability Engineer on telemetry requirements and with Data Engineering & Analytics on shared data-platform capabilities and integration patterns. Key Responsibilities Logging Platform Engineering Design, build, and operate logging capabilities across on-premise infrastructure, networks, applications, cloud platforms, and hybrid environments Collect logs from servers, network devices, applications, containers, databases, security appliances, cloud services, and other infrastructure sources Build scalable log ingestion, routing, enrichment, storage, indexing, search, and retrieval capabilities Define structured logging standards, schemas, metadata, tagging, and correlation conventions across services Design appropriate retention, archival, lifecycle, and deletion policies for different classes of operational data Support correlation between logs, metrics, events, and traces using common identifiers and telemetry standards Work with the Observability Engineer to implement platform capabilities supporting end-to-end service observability Data Collection & Integration Design secure and resilient data movement between on-premise environments, GCC, and approved external services Implement collection and forwarding patterns appropriate to different infrastructure, application, network, and security environments Design for intermittent connectivity, network constraints, buffering, retry, back-pressure, and recovery between environments Build event-driven and streaming patterns for moving operational data between producers and consumers Integrate legacy and enterprise systems with modern cloud-native platform capabilities Define clear interfaces and integration patterns between logging, observability, data engineering, and application platforms Data Platform Engineering Build shared platform capabilities for ingesting, storing, processing, querying, and serving operational data Design scalable storage and query architectures appropriate to data volume, access patterns, retention requirements, and cost Build ingestion, filtering, enrichment, and transformation pipelines for operational data Provide APIs, query interfaces, or other serving mechanisms for authorised downstream consumers Support operational datasets consumed by the User Portal, dashboards, reporting, automation, and Data Engineering & Analytics Define schemas and data contracts for shared platform interfaces Ensure platform changes remain backwards compatible or are coordinated with downstream consumers Cloud & Platform Engineering Design solutions using cloud-native logging, streaming, storage, search, and data capabilities Build infrastructure and platform configuration using Infrastructure as Code Automate build, test, deployment, configuration, and platform changes through CI/CD Design for scalability, resilience, high availability, recoverability, and operational simplicity Monitor platform capacity, performance, reliability, and cost Security & Governance Enforce MOE and Government data-classification requirements Design secure routing and storage of operational data across security zones and environment boundaries Apply appropriate encryption, access controls, authentication, and authorisation Ensure logging pipelines do not unnecessarily expose credentials, secrets, or sensitive information Implement audit-trail preservation and appropriate retention controls Ensure data-residency requirements are considered when routing operational data between on-premise, GCC, cloud, and SaaS environments Participate in security, architecture, and operational-readiness reviews Reliability & Operations Define SLOs and operational health indicators for logging and data-platform services Build monitoring, alerting, failure detection, retry, and recovery into platform components Monitor ingestion health, processing latency, data loss, storage utilisation, search performance, and platform availability Participate in incident investigation, root-cause analysis, and post-incident reviews Participate in operational support and on-call responsibilities for owned services Maintain architecture documentation, operational procedures, and runbooks What We Are Looking For Experience Minimum 3–5 years of experience in cloud engineering, platform engineering, DevOps, SRE, logging engineering, data platform engineering, or a related discipline At least 2 years of hands-on experience building or operating production logging, telemetry, or data-platform capabilities Demonstrated experience working with AWS and/or Azure cloud-native services Experience integrating on-premise and cloud environments, or operating systems in a hybrid environment Experience with production data ingestion, streaming, routing, storage, indexing, or search platforms Experience implementing Infrastructure as Code and CI/CD for production environments Experience designing systems for scalability, resilience, security, and operational support

Overview

Overview

We are looking for a Logging & Data Platform Engineer to design, build, and operate the logging and operational data-platform capabilities supporting the future SSOE platform.

You will build the platform that collects, transports, stores, indexes, searches, and serves operational data across MOE's technology environment, spanning on-premise infrastructure, networks, applications, GCC, AWS, Azure, and hybrid environments.

What You Will Be Working On

What You Will Be Working On
  • As a Logging & Data Platform Engineer, you will build the shared platform capabilities that enable engineering and operations teams to reliably collect and use operational data at scale.
  • You will work across logging, telemetry ingestion, data movement, storage, search, retention, and platform integration.
  • The role requires an engineer who understands both traditional enterprise infrastructure and modern cloud-native architectures and can design solutions that work securely and reliably across environment boundaries.
  • You will work closely with the Observability Engineer on telemetry requirements and with Data Engineering & Analytics on shared data-platform capabilities and integration patterns.

Key Responsibilities

Key Responsibilities

Logging Platform Engineering

  • Design, build, and operate logging capabilities across on-premise infrastructure, networks, applications, cloud platforms, and hybrid environments
  • Collect logs from servers, network devices, applications, containers, databases, security appliances, cloud services, and other infrastructure sources
  • Build scalable log ingestion, routing, enrichment, storage, indexing, search, and retrieval capabilities
  • Define structured logging standards, schemas, metadata, tagging, and correlation conventions across services
  • Design appropriate retention, archival, lifecycle, and deletion policies for different classes of operational data
  • Support correlation between logs, metrics, events, and traces using common identifiers and telemetry standards
  • Work with the Observability Engineer to implement platform capabilities supporting end-to-end service observability

Data Collection & Integration

  • Design secure and resilient data movement between on-premise environments, GCC, and approved external services
  • Implement collection and forwarding patterns appropriate to different infrastructure, application, network, and security environments
  • Design for intermittent connectivity, network constraints, buffering, retry, back-pressure, and recovery between environments
  • Build event-driven and streaming patterns for moving operational data between producers and consumers
  • Integrate legacy and enterprise systems with modern cloud-native platform capabilities
  • Define clear interfaces and integration patterns between logging, observability, data engineering, and application platforms

Data Platform Engineering

  • Build shared platform capabilities for ingesting, storing, processing, querying, and serving operational data
  • Design scalable storage and query architectures appropriate to data volume, access patterns, retention requirements, and cost
  • Build ingestion, filtering, enrichment, and transformation pipelines for operational data
  • Provide APIs, query interfaces, or other serving mechanisms for authorised downstream consumers
  • Support operational datasets consumed by the User Portal, dashboards, reporting, automation, and Data Engineering & Analytics
  • Define schemas and data contracts for shared platform interfaces
  • Ensure platform changes remain backwards compatible or are coordinated with downstream consumers

Cloud & Platform Engineering

  • Design solutions using cloud-native logging, streaming, storage, search, and data capabilities
  • Build infrastructure and platform configuration using Infrastructure as Code
  • Automate build, test, deployment, configuration, and platform changes through CI/CD
  • Design for scalability, resilience, high availability, recoverability, and operational simplicity
  • Monitor platform capacity, performance, reliability, and cost

Security & Governance

  • Enforce MOE and Government data-classification requirements
  • Design secure routing and storage of operational data across security zones and environment boundaries
  • Apply appropriate encryption, access controls, authentication, and authorisation
  • Ensure logging pipelines do not unnecessarily expose credentials, secrets, or sensitive information
  • Implement audit-trail preservation and appropriate retention controls
  • Ensure data-residency requirements are considered when routing operational data between on-premise, GCC, cloud, and SaaS environments
  • Participate in security, architecture, and operational-readiness reviews

Reliability & Operations

  • Define SLOs and operational health indicators for logging and data-platform services
  • Build monitoring, alerting, failure detection, retry, and recovery into platform components
  • Monitor ingestion health, processing latency, data loss, storage utilisation, search performance, and platform availability
  • Participate in incident investigation, root-cause analysis, and post-incident reviews
  • Participate in operational support and on-call responsibilities for owned services
  • Maintain architecture documentation, operational procedures, and runbooks

What We Are Looking For

What We Are Looking For

Experience

  • Minimum 3–5 years of experience in cloud engineering, platform engineering, DevOps, SRE, logging engineering, data platform engineering, or a related discipline
  • At least 2 years of hands-on experience building or operating production logging, telemetry, or data-platform capabilities
  • Demonstrated experience working with AWS and/or Azure cloud-native services
  • Experience integrating on-premise and cloud environments, or operating systems in a hybrid environment
  • Experience with production data ingestion, streaming, routing, storage, indexing, or search platforms
  • Experience implementing Infrastructure as Code and CI/CD for production environments
  • Experience designing systems for scalability, resilience, security, and operational support