Free with a WorkJio account: we compare your profile with this job and show what you already meet.
Share with a friend
Start
Not stated
Work mode
On-site
Experience
Experienced
About the role
This role owns the enterprise observability and log analytics platforms for a large organisation in Singapore, covering log search, APM, RUM and infrastructure monitoring across a hybrid-cloud estate. You will work with many application teams to roll out telemetry, keep costs and cardinality in check, and run the platforms as code. It suits an experienced platform or SRE engineer who has been the owner of an observability platform rather than only a user.
What you'll do
Administer the log and search platform, covering indexes, data onboarding, roles and access, apps and knowledge objects, retention, and search and dashboard performance.
Watch ingest latency, skipped searches and licence usage, and tune the platform to control cost.
Administer the APM, RUM and infrastructure monitoring platform, including teams, tokens, integrations, detectors, alert routing and dashboards.
Manage metric cardinality and usage against entitlement.
Roll out telemetry to all IS applications using OpenTelemetry Collector on AWS EKS and EC2, plus APM instrumentation for Java, Go, .NET, Node.js and Python.
Keep instrumentation vendor-neutral where possible so the backend can change without re-instrumenting, maintain onboarding guides, and report coverage per application.
Set up RUM for customer-facing web and mobile frontends and link browser sessions to backend traces, logs and infrastructure metrics.
Build and maintain observability as code in Git with peer review and CI/CD so detectors, dashboards and configuration are versioned and repeatable.
Plan and run platform and agent upgrades with vendors, including test plans, non-production validation, sign-off before release, and support cases and escalations.
Run structured evaluations and proofs of concept when a new or replacement tool is considered.
Design, build and operate AI and agent-assisted observability workflows such as alert noise reduction, automated triage, root-cause summaries and natural-language querying, with human review and guardrails.
Apply controls aligned with MAS-TRM and CSA best practices to telemetry, covering retention, access control, audit logging, and masking of PII and sensitive data before ingest.
Perform User Access Reviews on the observability platforms at least twice a year, covering accounts, roles and tokens, and follow up on removals and exceptions.
Take part in company-wide audits when required, providing evidence and closing findings on time.
Good to know
Certifications such as Splunk Core Certified Admin or Power User, Splunk Observability Cloud certification, or equivalent vendor certifications on another observability platform are a plus.
Requirements
Bachelor's degree in Computer Science, Information Technology, Engineering or a related field
5+ years in observability, monitoring, SRE or platform engineering
Hands-on administration, optimisation and health monitoring of an enterprise observability or log analytics platform
Splunk Cloud or Enterprise experience strongly preferred
Hands-on experience with an enterprise APM, RUM and infrastructure monitoring platform
Understanding of distributed tracing from frontend to backend
Apply on the company website
You will leave WorkJio. Never pay a fee or share your Singpass password, OTP or bank log-ins to apply.