Design,
develop, and maintain prompt frameworks, including system prompts,
few-shot examples, role-based prompts, and reasoning workflows for
production-grade LLM applications.
Build and
manage automated evaluation frameworks to measure model performance,
accuracy, latency, and regression across releases.
Conduct
structured A/B testing across prompt variations, model versions, and
configuration settings to optimize task-specific outcomes.
Convert
product requirements and edge-case scenarios into effective prompt
instructions, personas, constraints, and guardrails.
Partner
with ML engineers and product teams to determine when prompt engineering
is sufficient versus when fine-tuning, RAG, or other AI architectures are
required.
Create and
maintain a centralized prompt repository with version control, documentation,
and performance benchmarks for organizational reuse.
Lead
red-teaming and adversarial testing exercises to identify jailbreak risks,
hallucinations, and model vulnerabilities.
Define
evaluation criteria, annotation guidelines, and quality standards to
ensure consistency, safety, and reliability of AI-generated outputs.
Mentor
engineers and stakeholders on prompt engineering best practices,
evaluation methodologies, and the capabilities and limitations of modern
LLMs.
Present
prompt strategies, benchmark results, and trade-off analyses to product,
engineering, and leadership teams.
Apply
advanced prompting techniques, including chain-of-thought, zero-shot,
few-shot, and role-based prompting.
Drive
prompt testing, evaluation, benchmarking, and continuous optimization
efforts.
Improve AI
response quality through systematic assessment, tuning, and refinement.
Manage
context handling and prompt orchestration for complex AI workflows.
Technical Skills
Strong
programming and scripting skills in one or more modern programming
languages(C#, Python, Javascript).
Experience
building automation, evaluation pipelines, APIs, or AI-powered
applications using enterprise-grade development practices.
Hands-on
experience with LLM platforms, prompt engineering, model evaluation, and
AI application development.
Familiarity
with prompt orchestration frameworks, vector databases, RAG architectures,
and AI agent workflows.
Understanding
of data analysis, experimentation, benchmarking, and performance
optimization.
Experience
with version control systems, CI/CD pipelines, and cloud platforms.
Strong
knowledge of REST APIs, JSON, and system integration patterns.
Ability to
collaborate effectively with software engineers, data scientists, and
product teams to deliver production-ready AI solutions.
Experience Requirements
4-7 years
of combined experience in NLP, AI/ML products, software development,
technical writing, or related fields.
At least 2
years of direct, hands-on prompt engineering experience with production
LLM applications.
Proven
track record of owning and managing prompt systems end-to-end, from design
and implementation through monitoring and optimization in production.