Machine Learning Engineer - 2

SatSure
SatSure

Software Engineering

Posted on Jul 24, 2026
SatSure is a global Earth intelligence company headquartered in India. Founded in 2017, SatSure owns the full Earth observation data value chain – from the upstream payload infrastructure to the foundational deep-tech and AI layers, to the downstream decision intelligence solutions. We work with private companies and government bodies across Agri-tech, Agri-banking, forestry, critical infrastructure, and aviation. Through our subsidiary KaleidEO, SatSure is building multispectral payloads for high-definition space imagery with edge computing capabilities, powering both sovereign and commercial applications.
Role:
We are looking for a Machine Learning Engineer II to own a model workstream end-to-end — not a single model, but the family of models behind a product line, and the decisions that keep them accurate, fast, and affordable in production. You will be handed goals and ideas, not finished solutions: you decide how to run the experiments to test them, own the optimization and serving strategy, and defend your choices with benchmarks and clear findings. This is a hands-on role. You will still write the training loop and read the Triton logs.
Responsibilities:
  • Experiment Execution & Reporting: Own how a workstream's ideas get tested. Take the team's modelling proposals, turn them into well-run, reproducible training/fine-tuning experiments, monitor them, and report rigorous findings that decide what ships. You make pragmatic implementation choices; the research direction comes from Data Science.
  • Model Optimization: Own the accuracy/latency/cost tradeoff for your models in production. Drive quantization (INT8/FP16), pruning, distillation, and ONNX/TensorRT export; profile GPU utilization and eliminate the bottleneck rather than guessing at it.
  • Pipelines at Scale: Design ML pipelines that survive real data — petabyte-scale satellite archives, missing tiles, sensor drift, inconsistent projections. Make them reproducible and cheap to re-run.
  • Productionization: Prepare models for production and support their deployment on KServe / Triton alongside the Platform team. You should understand how modern serving frameworks work and what they demand of a model — batch sizing, concurrency, autoscaling behaviour, and failure modes under load — and hand over artifacts that account for them.
  • Evaluation & Benchmarking: Define what "good" means before training starts — offline metrics, production SLOs, and the acceptance criteria a model must meet to ship.
  • Monitoring & Drift: Own post-deployment model health. Instrument for drift, set thresholds, and drive the retraining decision rather than waiting to be told.
  • Technical Mentorship: Review code, models, and experiment design from MLEs/Data Scientists. Raise the floor of the team's engineering practice.
Qualification:
  • 4–7 years of relevant experience as a Machine Learning Engineer or in an applied research/engineering role.
  • Mandatory: Deep hands-on PyTorch expertise — custom datasets, distributed/multi-GPU training, mixed precision, and inference optimization. You can read a PyTorch profiler trace and act on it.
  • Mandatory: Multiple models shipped to production and kept in production, with evidence of measured latency/throughput improvements.
  • Mandatory: Hands-on experience training, fine-tuning, debugging, optimizing, and productionizing modern deep-learning architectures, including CNNs, vision transformers, and vision foundation models such as DINO- and SAM-style models. You are comfortable reading unfamiliar model implementations, adapting them to new use cases, and improving their training efficiency, inference performance, and production reliability.
  • Mandatory: Experience running experiments to a defined hypothesis with minimal supervision — and reporting findings others could act on.
  • Bachelor's degree in Computer Science, IT, Statistics, or a related field; non-IT degrees with strong relevant experience are acceptable.
Must-have skills:
  • ML Engineering Depth: You have debugged the training instability, found the data leak, and traced the production regression back to a preprocessing change. You know where models break in the real world.
  • Training at Scale: Confident running distributed / multi-GPU training and fine-tuning jobs efficiently and reproducibly, and instrumenting them so the findings are trustworthy.
  • Optimization Fluency: Quantization, distillation, ONNX/TensorRT, and batching/concurrency tuning at the serving layer. You can quantify what each technique bought you.
  • Feature & Data Pipelines: Scalable pipelines over raster/vector geospatial formats; comfortable with Rasterio/GDAL, tiling strategies, and the failure modes of remote sensing data.
  • Python & Engineering Practice: Clean, tested, maintainable code. You write ML code other engineers can pick up.
  • Containerization & Cloud: Confident with Docker and AWS (S3, EC2, ECR); you can reason about GPU instance selection and the cost of your own training runs.
  • Serving Infrastructure: Familiarity with KServe, Triton, or an equivalent model-serving stack — enough to understand how your model will be served and to work effectively with the team that runs it.
  • Experimentation & Versioning: MLflow (or equivalent) used properly — reproducible experiments, a model registry others can trust, and lineage from data to deployed artifact.
  • Kubernetes (User Level): Submitting GPU jobs, reading pod logs, reasoning about resource requests and limits.
Good-to-have:
  • Background in geospatial or remote sensing ML (satellite imagery, SAR, multispectral, time-series of Earth observation data).
  • Kernel-level bottleneck analysis and the ability to quantify what it bought you.
  • CI/CD for ML — automated training triggers, evaluation gates, and progressive rollout of models.
  • CUDA-level debugging and custom kernel awareness.
  • Experience with distributed training frameworks and large-scale data loading optimization.
  • Exposure to cost optimization for GPU workloads (spot strategy, right-sizing, inference cost per prediction).
Competencies:
  • Ownership: You own the outcome, not the artifact. If the model is slow in production, you drive the diagnosis and the fix with whoever owns the infrastructure it runs on.
  • Judgement: You know how to run an idea cheaply enough to get a signal fast, and when a negative result is conclusive. You do not over-engineer.
  • Scientific Rigor: Hypothesis-driven experimentation with results others can reproduce from your tracking and write-ups alone.
  • Collaboration & Influence: You can explain a tradeoff to a Data Scientist, a Platform engineer, and a domain expert — and get all three to agree on a path.
  • Mentorship: You make the engineers around you better through review, pairing, and clear technical writing.
Interview Process:
  • Intro call
  • Take-home assessment (focus on model training, optimization & deployment)
  • Interview rounds (ideally up to 3 rounds, including a deep dive on a model you have shipped and an ML system design discussion)
  • Culture round / HR round