Role summary
This work includes multi-agent orchestration, guardrails, agent runtime security, RAG, tool use, model customization, policy enforcement, OpenShell-like execution environments, and confidential AI deployments on protected infrastructure.
Responsibilities
- Lead strategic agentic AI partner engagements from discovery and architecture through PoC, production readiness, rollout, and scale.
- Build enterprise-grade agentic AI systems with multi-agent workflows, tool-using agents, RAG, planning, memory, evaluation, guardrails, policy enforcement, and failure containment.
- Partner with security ISVs to integrate the Company models into detection and response products, including threat triage, investigation agents, remediation workflows, natural-language-to-query, analyst automation, PII handling, and content safety.
- Architect secure and confidential AI deployments using the Company Confidential Computing, GPU attestation, KMS integration, protected infrastructure, air-gapped patterns, and partner key-management workflows.
- Create PoCs, benchmarks, reference architectures, reusable blueprints, field guidance, and product feedback that help the Company and our partners move secure AI systems into production.
Requirements
- BS, MS, or PhD in Computer Science, Electrical Engineering, AI/ML, or equivalent experience
- 8+ years in engineering, solutions architecture, applied ML, enterprise software, or technical deployment.
- Experience leading AI, ML, distributed systems, or enterprise software projects from prototype to production.
- Hands-on experience building LLM, generative AI, RAG, or agentic AI applications in production or production-like environments.
- Depth in one or more areas such as AI/LLM security, enterprise cybersecurity, trust and safety, confidential computing, secure AI infrastructure, model customization, post-training, or model evaluation.
- Strong Python and Linux skills, experience with PyTorch, TensorFlow, or similar frameworks, and working knowledge of risks such as prompt injection, jailbreaks, tool-based data exfiltration, unsafe tool invocation, and model or skill supply-chain risk.
Nice to have
- Experience with the Company AI software such as NIM, NeMo Framework, NeMo Retriever, NeMo Guardrails, NeMo Agent Toolkit, Dynamo, Nemotron, Nemotron Safety models, Triton, TensorRT-LLM, or NIM Operator.
- Experience with LLM red-teaming, AI safety evaluation, adversarial testing, prompt-injection defense, policy enforcement, Garak, NeMo Auditor, or release-gating evaluation benchmarks.
- Experience with OpenShell, agent harnesses, sandboxed execution, secure tool invocation, agent runtime security, AI/software supply-chain security, model or skill signing, provenance, attestation, VEX, or secure model registries.
- Experience building post-training pipelines or GPU-accelerated safety and security workflows, including reasoning, tool use, domain adaptation, safety alignment, PII/NER detection, content-safety models, TensorRT optimization, quantization, workshops, architecture reviews, whitepapers, or reference architectures.
- Experience with confidential computing, including GPU confidential computing, remote attestation, Confidential Containers, enterprise KMS, air-gapped deployments, AMD SEV-SNP, or Intel TDX.