Logo-of-Ignatiuz-hiring-for-jobs-in-India-on-GrabJobs

AI Infrastructure and Platform Architect

Job Description - AI Infrastructure and Platform Architect

Position Summary

We are seeking an experienced AI Infrastructure and
Platform Architect
to design, optimize, and manage scalable AI
infrastructure and platforms across on-premises, cloud, and hybrid
environments.

The ideal candidate should have strong experience with
GPU-based systems, AI/ML platforms, infrastructure architecture, performance
optimization, capacity planning, and production support. The role will work
closely with AI/ML developers, DevOps engineers, data engineers, and solution
architects to improve the performance, reliability, scalability, and cost
efficiency of AI solutions.

Key Responsibilities

  • Design
    and manage AI infrastructure for model training, fine-tuning, inference,
    computer vision, Generative AI, and LLM workloads.
  • Define
    CPU, GPU, RAM, VRAM, storage, networking, cooling, and power requirements.
  • Review
    existing hardware and platform performance and recommend upgrades or
    optimizations.
  • Perform
    capacity planning to support future workloads and minimize frequent
    hardware changes.
  • Build
    and maintain AI platforms using Linux, Docker, Kubernetes, GPU
    orchestration, and cloud services.
  • Configure
    and manage NVIDIA drivers, CUDA, cuDNN, TensorRT, and related AI
    acceleration technologies.
  • Monitor
    system health, GPU utilization, memory usage, storage performance, and
    network throughput.
  • Diagnose
    infrastructure failures, system crashes, performance bottlenecks, and
    platform outages.
  • Implement
    monitoring, alerting, backup, disaster recovery, security, and operational
    best practices.
  • Prepare
    architecture documents, hardware specifications, technical
    recommendations, and operational runbooks.
  • Support
    production deployment, troubleshooting, and continuous platform
    improvement.

AI Solution Optimization

The candidate should also be capable of:

  • Reviewing
    the end-to-end AI solution and identifying performance, architecture, and
    infrastructure gaps.
  • Recommending
    improvements to scalability, reliability, maintainability, and cost
    efficiency.
  • Supporting
    AI/ML developers with model training and experimentation environments.
  • Helping
    reduce training time through GPU optimization, distributed training,
    resource tuning, and efficient data pipelines.
  • Providing
    guidance on model accuracy, evaluation, hyperparameter tuning, and
    experimentation practices.
  • Improving
    model-serving and inference performance.
  • Mentoring
    existing team members on AI infrastructure and production-readiness best
    practices.

Required Skills

  • AI
    infrastructure and GPU-based computing
  • NVIDIA
    GPU architecture, CUDA, cuDNN, NCCL, and TensorRT
  • Linux
    administration
  • Docker
    and Kubernetes
  • PyTorch,
    TensorFlow, Hugging Face, or similar frameworks
  • Cloud
    and on-premises AI platforms
  • Infrastructure
    sizing and capacity planning
  • Performance
    monitoring and troubleshooting
  • High-performance
    storage and networking
  • MLOps,
    CI/CD, automation, and Infrastructure as Code
  • Monitoring
    tools such as Prometheus, Grafana, NVIDIA DCGM, or OpenTelemetry

Qualifications

  • Bachelor's
    or Master's degree in Computer Science, Artificial Intelligence,
    Information Technology, Engineering, or a related field.
  • Strong
    overall experience in infrastructure, cloud, platform engineering,
    architecture, or AI systems.
  • A
    minimum of 3 years of direct, hands-on experience specifically working
    with AI infrastructure, machine learning platforms, GPU environments, or
    production AI workloads
    .
  • Proven
    experience designing, deploying, supporting, or optimizing production AI
    systems.
  • Strong
    problem-solving, troubleshooting, communication, and technical
    documentation skills.

The candidate's total professional experience may be
significantly higher. However, at least three years should involve genuine,
relevant, hands-on work with AI systems and platforms.

Added Advantage

Preference will be given to candidates who have experience
with:

  • AI
    solution architecture
  • Large
    Language Models and Generative AI
  • RAG
    and agentic AI systems
  • Distributed
    model training
  • Computer
    vision and edge AI
  • Model-serving
    platforms
  • AI
    performance benchmarking
  • FinOps
    and infrastructure cost optimization
  • High
    Performance Computing environments

Experience Validation

Candidates should be able to explain their direct
contribution to AI projects, including:

  • AI
    infrastructure or platforms they designed or managed
  • GPU
    and hardware-sizing decisions
  • Model
    training or inference environments supported
  • Performance
    issues diagnosed and resolved
  • Improvements
    achieved in training time, utilization, reliability, or cost
  • Production
    AI workloads they deployed or maintained

General DevOps, cloud, or system administration experience
without direct AI or machine learning exposure will not be sufficient for this
position.



Original job AI Infrastructure and Platform Architect posted on GrabJobs ©. To flag any issues with this job please use the Report Job button on GrabJobs.
Share Job
Share Job

Similar AI Infrastructure and Platform Architect Jobs in India

GrabJobs is the no1 job portal in India, connecting you to thousands of jobs fast! Find the best jobs in India, apply in 1 click and get a job today!

Mobile Apps

Copyright © 2026 Grabjobs Pte.Ltd. All Rights Reserved.