100+ Machine Operation Jobs - July 2026 - High Salaries

显示193个工作的结果 "machine operation"
不要错过任何 Machine Operation 的新工作机会
Undisclosed

KL City

  • Investigate issues, conduct initial issue triage, and support resolution of production model operations.
  • Document model behaviour, validation outcomes, testing results, and performance outputs.
  • Work with Python-based modelling and analytics solutions within Databricks environments. ...
Posted
16 days ago
Undisclosed

Singapore

  • Dive deep into machine learning methods, including large language models (LLMs), multimodal models, and image/video generation models, adapting these methods to TikTok Live scenarios and serving as foundational models and features.
  • Currently pursuing a Master's degree in Computer Science.
  • Familiarity with large language models and its applications ...
Posted
13 days ago
Undisclosed

KL City

  • Build and maintain CI/CD pipelines for machine learning solutions, enabling automated testing, packaging, deployment, and environment management.
  • Monitor production ML models and data pipelines, proactively addressing performance, data quality, model drift, latency, and operational issues.
  • Collaborate with data engineering, data science, security, and governance teams to deliver reliable, scalable, and compliant AI/ML solutions. ...
Posted
2 days ago
Undisclosed

Singapore

  • Leverage AWS services (SageMaker, ECR, EKS/ECS, Lambda, Step Functions, S3, CloudWatch, CloudFormation, Terraform, etc.) to host, scale, and manage model training and inference pipelines.
  • Develop monitoring and alerting solutions for model latency, accuracy, data drift, and infrastructure health; integrate with Prometheus, Grafana, CloudWatch, or similar tools.
  • Automate model versioning, artifact storage, and metadata tracking using Mlflow or SageMaker model registry. ...
Posted
13 days ago
Undisclosed

Singapore

  • Architect and implement solutions leveraging AWS services (SageMaker, ECR, EKS/ECS, Lambda, S3, CloudWatch, CloudFormation, Terraform, etc.) to host, scale, and manage model training and inference pipelines.
  • Develop comprehensive monitoring and alerting solutions for model latency, accuracy, data drift, and infrastructure health; integrate with Prometheus, Grafana, CloudWatch, or similar tools.
  • Oversee the automation of model versioning, artifact storage, and metadata tracking using Mlflow or SageMaker model registry. ...
Posted
18 days ago
Undisclosed

KL City

  • Collaborate closely with IT operations and DevOps teams to ensure the smooth integration of infrastructure platforms with other applications and processes.
  • Establish and maintain systems for monitoring machine learning models in production. Oversee the development of machine learning models at MoneyLion to ensure proper model governance.
  • Effectively manage cloud infrastructure costs by monitoring and optimizing spending, and provide transparency and accountability in cost-related matters. ...
Posted
9 days ago

CONSTRUCTOR TECHNOLOGY PTE. LTD.

SGD6,000 - SGD10,000 每月

Singapore

  • Implement monitoring,observability, and alerting for models in production.
  • Manage the model lifecycle withMLflow: experiment tracking, versioning, registries, and reproducibility.
  • Automate infrastructure andpartner with ML and platform teams on standards. ...
Posted
22 days ago

CONSTRUCTOR TECHNOLOGY PTE. LTD.

SGD6,000 - SGD6,000 每月

Singapore

  • ·      Design and maintain CI/CDpipelines for ML models and services.
  • ·      Build and operate modeldeployment, serving, and rollback workflows.
  • ·      Implement monitoring,observability, and alerting for models in production. ...
Posted
23 days ago
Undisclosed

Singapore

  • You will also be responsible for training stability and reliability. This includes identifying the root causes of loss spikes, divergence, slow nodes, communication bottlenecks, checkpoint failures, and data-related instability, as well as designing mechanisms for fast checkpoint recovery and automatic exclusion of problematic nodes.
  • The ideal candidate has strong hands-on experience with PyTorch distributed training and a solid understanding of CUDA architecture, GPU memory hierarchy, NCCL communication, and performance profiling.
  • You should have source-level familiarity with at least one major large-scale training framework, such as Megatron-LM, DeepSpeed, PyTorch FSDP, or TorchTitan, and be comfortable reading, modifying, and debugging framework internals. ...
Posted
24 days ago
Undisclosed

KL City

  • Assist Data Scientists with resource provisioning in the model development and deployment phase.
  • Help maintain monitoring systems for machine learning models in production environments.
  • Contribute to optimizing cloud infrastructure usage and enhancing transparency around resource costs. ...
Posted
a month ago
Undisclosed

最后机会申请此工作。

Posted
11 years ago
MYR1,200 - MYR800 每月

最后机会申请此工作。

Posted
10 years ago
高机会
MYR1,200 - MYR800 每月

最后机会申请此工作。

Posted
10 years ago