Czarodzieje.AI

T-Hub - AIOps Engineer - AI Infrastructure & Orchestration

T-Mobile Warszawa, Mokotów Mid

Wynagrodzenie do uzgodnienia

🪄 Prompt EngineeringStacjonarnieB2B CONTRACT

Aplikuj na tę ofertę

Wyślemy Twój profil bezpośrednio do firmy.

O roli

## Your responsibilities - Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure. - Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants. - Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3. - Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency. - Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing. - Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates. - Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements. - Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics. - Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns. - Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services. - Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices. - Maintain audit logging capabilities to support compliance, security investigations, and operational governance. - Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services. - Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives. - 5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations. - At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms. - Strong experience with Kubernetes and OpenShift administration in production environments. - Proven experience deploying and operating vLLM-based inference platforms in production. - Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques. - Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack. - Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms. - Strong Python programming skills with experience developing automation and operational tooling. - Experience with Bash scripting and Linux systems administration. - Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices. - Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills. ## What we offer - Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation. - You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity. - sharing the costs of sports activities - private medical care - sharing the costs of professional training & courses - life insurance - remote work opportunities - flexible working time - corporate products and services at discounted prices - mobile phone available for private use - no dress code - parking space for employees - extra social benefits - sharing the costs of tickets to the movies, theater - holiday funds - birthday celebration - sharing the costs of a streaming platform subscription - employee referral program - charity initiatives - extra leave - platforma benefitowa ## Recruitment stages - Prześlij swoje CV - Spotkaj się z przyszłym liderem/ liderką zespołu - Witaj w T-Mobile :) - We are a technology company, and our goal is to create innovative solutions for individual and business clients. - At T-Mobile, we all live in a magenta world! This color is close to our hearts and means faith in the success of undertaken actions, self-confidence, and endurance. - That’s who we are as a team. - At #MagentaTeam , we focus on exchanging experiences, agile work, and quick adaptation to changes! #MagentaTeam is, above all, a mix of different competencies, experiences, personalities, temperaments, and views. And this diversity is our greatest strength.

Obowiązki

Wymagania