Design, develop, and maintain prompt frameworks, including system prompts, few-shot examples, role-based prompts, and reasoning workflows for production-grade LLM applications.
Build and manage automated evaluation frameworks to measure model performance, accuracy, latency, and regression across releases.
Conduct structured A/B testing across prompt variations, model versions, and configuration settings to optimize task-specific outcomes.
Convert product requirements and edge-case scenarios into effective prompt instructions, personas, constraints, and guardrails.
Partner with ML engineers and product teams to determine when prompt engineering is sufficient versus when fine-tuning, RAG, or other AI architectures are required.
Create and maintain a centralized prompt repository with version control, documentation, and performance benchmarks for organizational reuse.
Lead red-teaming and adversarial testing exercises to identify jailbreak risks, hallucinations, and model vulnerabilities.
Define evaluation criteria, annotation guidelines, and quality standards to ensure consistency, safety, and reliability of AI-generated outputs.
Mentor engineers and stakeholders on prompt engineering best practices, evaluation methodologies, and the capabilities and limitations of modern LLMs.
Present prompt strategies, benchmark results, and trade-off analyses to product, engineering, and leadership teams.
Apply advanced prompting techniques, including chain-of-thought, zero-shot, few-shot, and role-based prompting.
Drive prompt testing, evaluation, benchmarking, and continuous optimization efforts.
Improve AI response quality through systematic assessment, tuning, and refinement.
Manage context handling and prompt orchestration for complex AI workflows.
Technical Skills
Strong programming and scripting skills in one or more modern programming languages(C#, Python, Javascript).
Experience building automation, evaluation pipelines, APIs, or AI-powered applications using enterprise-grade development practices.
Hands-on experience with LLM platforms, prompt engineering, model evaluation, and AI application development.
Familiarity with prompt orchestration frameworks, vector databases, RAG architectures, and AI agent workflows.
Understanding of data analysis, experimentation, benchmarking, and performance optimization.
Experience with version control systems, CI/CD pipelines, and cloud platforms.
Strong knowledge of REST APIs, JSON, and system integration patterns.
Ability to collaborate effectively with software engineers, data scientists, and product teams to deliver production-ready AI solutions.
Experience Requirements
4-7 years of combined experience in NLP, AI/ML products, software development, technical writing, or related fields.
At least 2 years of direct, hands-on prompt engineering experience with production LLM applications.
Proven track record of owning and managing prompt systems end-to-end, from design and implementation through monitoring and optimization in production.
All Job Ads are subject to GrabJobs’s Terms of Service. We allow users to flag postings that may be in violation of those terms. Job Ads may also be flagged by GrabJobs moderation team. However, no moderation system is perfect, and flagging a posting does not ensure that it will be removed.
Be the first to receive the latest Others Full-Time Jobs in Australia.
Setup your job alert:
By activating job alerts, I agree to GrabJobs Terms & Privacy Policy. I can unsubscribe to job alerts anytime.
Skip
GrabJobs is the no1 job portal in Australia, connecting you to thousands of jobs fast!
Find the best jobs in Australia, apply in 1 click and get a job today!