LI
Site Reliability Engineer
MaleField service engineerLive in SingaporeNationality
Share
Work experience
Site Reliability Engineer
TikTok2023.10-Current(3 years)• Ran water-level capacity governance for tier-0 services — continuous inspection with dynamic adjustment to each service’s healthy band — and onboarded and configured profile-based single-instance fault self-healing (automated instance migration+ load-balancer reweighting pull unhealthy instances within ~30s), preserving availability and burst headroom. • Drove a service-tier standardization program collapsing multi-cluster footprints into single-cluster deployments — cut operated physical clusters by 60% while lifting utilization, availability, and maintainability. Resource efficiency & FinOps • Built an internal resource-governance platform for the recommendation stack — utilization dashboards plus automated per-scenario governance workflows (e.g., low-utilization reclamation) — and ran it across online / offline compute (Flink /Spark), KV stores (Redis + in-house), and HDFS, reclaiming 100K+ cores and multiple petabytes reinvested into priority projects. • Built a full-stack ROI platform (React + Django) that computes project ROI from experiment-benefit and resource-cost data, directing limited resources to the highest-return projects. • Owned monthly cost attribution for live recommendation; rebuilt cross-region billing-push pipelines and drove pricing-policy reviews to balance cost across regions. New-capability qualification & rollout • Normalized colocation — onboarded offline-on-online-node colocation to absorb spare compute in low-utilization win- dows; gray-ramped without impacting online-link stability or latency. • Flexible-spec cluster migration (recommendation + search): – Qualify: new-architecture research, FMEA (failure mode & efects analysis), a test plan (coverage scope + performancetesting across service types), and multi-stage cluster testing (smoke, stress/limit, and long-duration observation at the same water level as production), with published reports. – Gray migration: set ramp cadence and service sequencing, observed stability through each phase, and raised memory utilization to progressively meet the assessment targets. Datacenter buildout & traffic migration • Supported repeated cross-region traffic scheduling and new-datacenter buildouts — deploying live-recommendation and application-compute link services, owning engineering- and business-metric acceptance, supporting experiment-metric par- ity, and troubleshooting anomalies. • Independently executed the live and live-commerce recommendation traffic migrations end-to-end — assessing link-service capacity, migrating services, monitoring engineering metrics, and resolving anomalies.System DevOps Engineer
SEA Group2022.06-2023.10(a year)• Built an internal infrastructure-provisioning platform end-to-end—aserver-provisioning ticketing module and a data-center parts-management module — and re-architected an ISP-management UI into a micro-frontend. • Migrated front- and back-end services onto a Rancher Kubernetes cluster — GitLab CI/CD (build/test/deploy), Harbor registry, dev/staging/prod environments, Nginx traffic cut-over, and monolith → microservices refactor.Software Engineer
Silver Factory Technology2021.09-2022.05(9 months)• Developed hardware-control applications for medical and gas-sensing devices (embedded C++, Python, Arduino, multi-threading, concurrency); owned the software test report for an ISO 13485-controlled COVID-19 detector.Equipment Engineer
UMC2018.04-2021.08(3 years)• Semiconductor equipment engineer; troubleshot tool failures and led two equipment cost-reduction projects (mechanicaltransfer-speed modification; quartz and parts lifetime extension).
Educational experience
Nanyang Technological University
of Mechanical Engineering2014.09-2017.11(3 years)
Languages
Chinese (Mandarin)
Native
English
Native
Resume Search
Nationality
Job category
City or country
Sort by
Contact way
65****1153
li**@**om
Membership will unlock the resume
Also view
