Job Overview
Job description
- Location:
- Dublin, Ireland
- Work arrangement:
- On-site
Site Reliability Engineer - Cloud Infrastructure at TikTok in Dublin, Ireland.
Role Summary
What the team does
- Reliability: Ensuring the reliability and efficiency of our core infrastructure, focusing on system capacity and stability; setting up reliability standards and recovery SOP.
- Reliability: Troubleshooting and locating the technical issues, bottleneck analysis, managing system high availability architecture transformation and upgrading.
- Efficiency: Building automated operation solutions for large-scale systems; partnering with system development teams for system iteration.
- Efficiency: Designing and implementing software platforms and monitoring frameworks for efficient, automated, and intelligent service-oriented architecture (SOA) governance.
- Cost: There are millions of CPUs. We should build delivery standards, and monitor and budget systems to optimize the cost of the company.
- Compliance: Designing and setting up new IDC; designing and implementing a data protection plan to meet the standard requirement.
Responsibilities
The Technical Infrastructure SRE team is responsible for managing the whole infrastructure and applications. Our mission is to ensure all production systems can support our fast growing world-wide user base as well as keep the entire systems stable, efficient and cost effective. We manage deployments, system capacity, traffic scheduling, fault tolerance, disaster recovery, emergency response, automations, operation platforms development, etc.
Our team is full of diversity. We have team members in Singapore and China. Now we are extending our teams to Ireland. We are looking forward to seeing new talents joining our team and together helping TikTok grow.
Be responsible for the basic engineering construction of byte infrastructure products & components, focusing on infrastructure O&M architecture optimization, automated O&M platform research and development, data and intelligent O&M. Through the methodology of software engineering and digital intelligence, O&M, around the O&M requirements of infrastructure products & components, built a layered and systematic O&M platform to solve the problem of ultra-large-scale cluster O&M management. (Goals) To provide stable, efficient, and low-cost serverless infrastructure facilities for Mid-Platform & Business. We aim to be the leading SRE team across the industry。
- Reliability: Ensure the stability of the company's core infrastructure (system high availability and reliability), focus on system performance and capacity, establish O&M (Operation & Maintenance) standards and SOP processes.
- Reliability: Troubleshooting and locating technical issues, collaborate with the technical team to develop and implement system capacity planning, performance testing, anomaly analysis, and fault diagnosis and resolution strategies.
- Efficiency: Research and evaluate large-scale system architectures and technologies, use new tools and technologies to improve existing systems and processes to support business development.
- Efficiency: Design and implement O&M platforms to achieve efficient, automated, and intelligent system maintenance.
- Cost: Develop delivery standards for mass production system scales, from budgeting to resource delivery, to online system capacity assessments, to help the company optimize IT costs.
- Compliance: Design and establish new IDC, design and implement data protection plans to meet standard requirements.
Requirements
- Bachelor's / Master's Degree in Computer Science or related major.
- Solid basic knowledge of computer software, understanding of Linux operating system, storage, network IO and other related principles.
- Familiar with one or more programming languages, such as Python, Go, and Java. Knowledge of design patterns and coding principles is necessary.
- Familiar with Elastic Search
Preferred Qualifications
- Experience with storage, and relevant system experience with the following: KV, Table, Graph, Redis, MySQL, MongoDB, MQ, and Kafka.
- Experience with computing & big data, and system experience with the following: Kubernetes, Docker/Containers, AIops, Spark, Flink, Function as a service, RPC Framework, and Service Mesh.
About the Company
The Global Business Solutions (GBS) team is responsible for the revenue growth of the TikTok business, and our teams include Sales, Marketing, Ops, Account Managers, Agency and partnerships, as well as Marketing Science.
At TikTok, our Global Business Solutions (GBS) team plays a key role in generating revenue by promoting our advertising solutions, onboarding new clients, driving ad campaigns, and more. As the TikTok community grows at an unprecedented speed around the world, our GBS team leads groundbreaking projects that are changing the landscape of the advertising industry in real time.
We're seeking an analytically driven, and detail-oriented Client Solutions Manager (CSM) Intern to join our Ecommerce Team. As a CSM, you will partner closely with Client Partners and Client Solutions Managers to drive revenue by identifying opportunities, leveraging data insights, and delivering consultative solutions for advertisers.
This role centers on client education, relationship growth, data analysis, and campaign success. You will provide strategic recommendations to both clients and internal teams, ensuring campaigns achieve business objectives while optimizing long-term partnerships. Success in this role requires strong data analytics skills, adaptability in a fast-paced environment, and a test-and-learn mindset to uncover the best solutions.
- Role:
- Site Reliability Engineer - Cloud Infrastructure
- Job Type:
- Mid Level | Software Engineering