Role and Responsibilities
Operational Responsibilities:
- Monitor and maintain production data pipelines to ensure 99.9% uptime and optimal performance
- Implement comprehensive logging, alerting, and monitoring systems using Application monitoring tools
- Perform regular health checks performance, job execution times, and resource utilization to identify and resolve bottlenecks proactively
- Manage incident response procedures for pipeline failures, including root cause analysis, resolution, and post-incident reviews
- Establish and maintain disaster recovery procedures and backup strategies for critical data assets within the Databricks environment
- Conduct regular performance tuning of Spark jobs and Databricks cluster configurations to optimize cost and execution efficiency
- Maintain comprehensive documentation for operational procedures, runbooks, and troubleshooting guides
- Coordinate scheduled maintenance windows and system upgrades with minimal business impact
- Manage user access controls, workspace configurations, and security policies within Application environments
Requirements / Qualifications
Education & Experience:
- Degree in Computer Science or Computer Engineering
- Minimum 5 years working experience in system operations compliance and management areas
- Project hands-on experience specifically with AWS platform (primary requirement)
- Project experience in cloud operations or cloud architecture
- Must be cloud certified (AWS)
Core Technical Skills:
- Proficiency in Databricks platform, including workspace management, cluster configuration, and job orchestration
- Strong expertise in Apache Spark within Databricks environment, including Spark SQL, DataFrames, and RDDs
- Good in-depth understanding of data warehouse concepts, data profiling, data verification and advanced analytics techniques
- Strong knowledge of monitoring, incident management, and cloud cost control
Technology Stack Experience:
- Databricks
- AWS cloud services and architecture
- IDMC (Informatica Data Management Cloud)
- Tableau for data visualization
- Oracle Database management
- ML Ops practices within Databricks environment
- STATA for statistical analysis is advantage
- Amazon SageMaker integration with Databricks
- DataRobot platform integration
Soft Skills & Stakeholder Management:
- Good interpersonal skills with the ability to work with different groups of stakeholders
- Strong problem-solving skills and ability to work independently in a fast-paced environment with minimal supervision
- Excellent communication skills for technical documentation and cross team collaboration
Desirable Requirements:
- AWS certification (Associate or Professional level) - highly preferred • Exposure to hospital information/clinical systems is an added advantage
- Understanding of DevOps practices and CI/CD pipelines for Databricks based data engineering projects
NTT Singapore Pte Ltd (NTTS)
is the regional headquarters of NTT Communications Corporation (NTT Com) for Asia Pacific Region.
Established in 1997, NTT Singapore has more than 10 years of expertise in providing information and communications technology (ICT) solutions worldwide.
NTT Singapore offers diverse high-quality connectivity, data centre solutions, security services, IT management services, voice and conferencing solutions and solution integration services to its enterprise customers.
NTT Communications is a wholly owned subsidiary of Nippon Telegraph and Telephone Corporation (NTT Corp.), one of the world’s largest providers of telecommunications services.
In 2013, NTT Corp. is ranked no.1 in telecom industry in the Fortune Global 500* list with operating revenues of more than $133,077 million. It is positioned 32nd among the top 500 corporations worldwide.
NTT Com's extensive global infrastructure includes Arcstar secure private networks, which cover 196 countries/regions and a tier-1 IP backbone network connected with major ISPs worldwide, as well as secure data centers at over 150 locations worldwide.