Description
Key Responsibilities
AI/ML Model Development
- Design and implement ML models for anomaly detection, predictive failure analysis, and connector health monitoring.
- Build and deploy supervised and unsupervised learning pipelines for IT operations analytics and AIOps use cases.
- Develop time-series forecasting models to anticipate connector degradation and L4 incident spikes.
- Implement model versioning, A/B testing, automated retraining, and drift monitoring pipelines.
- Maintain feature stores, data quality standards, and model registries aligned with MLOps best practices. LLM & Generative AI Integration
- Integrate LLM APIs such as AWS Bedrock, OpenAI, and Anthropic Claude into Concierto connector orchestration workflows.
- Build RAG pipelines for intelligent connector documentation search, incident summarization, and self-healing runbooks.
- Design and optimize prompt engineering strategies for root cause analysis, change advisory drafting, and test case generation.
- Develop AI agents using LangChain, CrewAI, AutoGen, or equivalent frameworks for autonomous incident triage and connector lifecycle management.
- Deploy and manage LLM inference endpoints on AWS Lambda, ECS, or SageMaker with IAM-secured access controls.
NLP & Intelligent Log Analytics
- Develop NLP-based log parsing, event correlation, and semantic classification modules to accelerate L4 support triage.
- Build natural language query interfaces for operations teams to interrogate connector telemetry and CloudWatch logs.
- Apply NER, intent classification, and text summarization to convert raw incident data into actionable insights.
- Implement vector search and semantic similarity using AWS OpenSearch, Pinecone, or equivalent solutions.
AI-Powered Automation Engineering
- Design agentic AI pipelines that autonomously diagnose, escalate, or resolve common Concierto connector issues.
- Build AI-augmented test automation frameworks that generate, execute, and evaluate test cases for connector APIs.
- Develop self-healing test scripts using pattern recognition and element-level ML locators.
- Integrate AI-based defect prediction and test coverage analysis into CI/CD pipelines.
- Automate connector deployment health checks, rollback triggers, and post-deployment validation using AI-driven observability.
AWS Cloud AI Integration
- Leverage AWS AI/ML services including SageMaker, Bedrock, Comprehend, Forecast, and OpenSearch.
- Integrate AI inference outputs with Java-based connector REST APIs through well-defined, versioned service contracts.
- Optimize ML model latency and throughput for real-time connector event classification and response at scale.
- Apply AWS security best practices including IAM, KMS, and VPC across AI data pipelines and inference endpoints.
- Design feature engineering pipelines from structured and semi-structured AWS event streams and connector logs.
Platform Observability & Intelligent Reporting
- Build AI-powered dashboards and alerts for Concierto connector KPIs using CloudWatch, Grafana, or equivalent tools.
- Generate AI-authored incident summaries, root cause analysis reports, and resolution recommendations.
- Develop executive-ready AI-generated performance narratives and weekly connector health digests.
Collaboration, Agile & Mentorship
- Work cross-functionally with Java, QA, DevOps, and Product Management teams to align AI modules with the Concierto Agentic roadmap.
- Participate in agile ceremonies including sprint planning, backlog grooming, retrospectives, and release planning.
- Maintain model cards, prompt libraries, API integration guides, and AI runbooks as living documentation assets.
- Mentor team members on AI/ML integration patterns, prompt engineering practices, and responsible AI principles.
Required Qualifications
Education & Experience
- Bachelor's or Master's degree in Computer Science, Data Science, AI/ML, Software Engineering, or equivalent.
- 4 to 7 years of hands-on experience in ML engineering, AI application development, or data science roles.
- Minimum 2 years of demonstrated experience integrating AI/LLM/NLP capabilities in production environments.
- Prior experience delivering cloud-native AI solutions on AWS. AI/ML & Data Science Skills
- Proficient in Python with ML frameworks such as scikit-learn, TensorFlow, PyTorch, or equivalent.
- Strong understanding of supervised and unsupervised learning, time-series forecasting, classification, and clustering.
- Hands-on experience with AWS SageMaker for model training, hosting, monitoring, and MLOps pipelines.
- Familiarity with MLflow, Kubeflow, or equivalent MLOps tooling.
- Proficient in feature engineering, data wrangling, and working with structured and unstructured data at scale. LLM, GenAI & NLP Skills
- Hands-on experience with LLM APIs such as OpenAI, AWS Bedrock, Anthropic, or Hugging Face.
- Practical knowledge of RAG architecture, vector databases, and embedding pipelines.
- Proficient in prompt engineering including zero-shot, few-shot, chain-of-thought, and tool-use patterns.
- Experience building NLP pipelines including tokenization, NER, intent classification, summarization, and sentiment analysis.
- Familiarity with agentic frameworks such as LangChain, CrewAI, AutoGen, or equivalent.
Software Engineering & AWS Integration
- Proficient in Python and/or Java with experience exposing and consuming REST APIs in microservices architectures.
- Hands-on experience with AWS services including Lambda, S3, SQS, SNS, CloudWatch, ECS, and IAM.
- Familiarity with CI/CD pipelines using Jenkins, GitHub Actions, or GitLab CI.
- Working knowledge of Docker, Kubernetes, or serverless deployment patterns for AI workloads.
- Proficient in SQL; experience with data lakes, streaming data, or event-driven architectures is an added advantage.
Preferred Qualifications
- AWS Certified Machine Learning - Specialty, AWS AI Practitioner, or AWS Certified Developer - Associate.
- Experience with enterprise identity and access management platforms or IT connector ecosystems such as IAM, OAuth 2.0, SCIM, or LDAP.
- Background in AIOps, ITSM automation, cybersecurity analytics, or cloud infrastructure intelligence.
- Exposure to responsible AI practices including bias detection, model explainability, SHAP/LIME, and AI governance frameworks.
- Experience with AI-enhanced test automation frameworks such as Selenium, Playwright, or Karate.
- Contributions to open-source AI/ML projects or published technical content on AI engineering topics