
NLP QA Engineer
BruntWork · Posted today
- India (Remote)
- Remote
- Contract
- 3+ yrs
- ₹15.6L
About the role
About the Role
We are seeking a detail-oriented NLP QA Engineer to join our technical team in ensuring the reliability, accuracy, and security of Candid Chat AI across both web and mobile platforms.
In this role, you will go beyond traditional software testing by actively evaluating Large Language Models (LLMs), conversational flows, and system integrations. You will design, execute, and automate tests to benchmark AI output quality, conduct rigorous adversarial and safety testing, and integrate automated quality gates into our CI/CD pipelines.
Schedule
- Monday to Friday, flexible schedule within the client’s business hours (Australian Eastern Time, 40 hours per week)
Key Responsibilities
- AI & Conversational Testing: Test end-to-end conversational AI flows across web and mobile interfaces to ensure seamless user experiences.
- Model Evaluation & Quality Assurance: Systematically evaluate AI responses for factual accuracy, context relevance, tone consistency, and hallucination detection.
- Adversarial & Safety Testing: Perform red-teaming and adversarial testing, including prompt injection, jailbreaking attempts, and guardrail verification to ensure model safety and compliance.
- Regression Testing: Design and execute regression test suites whenever model versions, system prompts, or backend configurations are updated.
- Test Automation: Build and maintain UI and API automated testing frameworks for frontend applications and underlying backend services.
- Benchmarking & Datasets: Construct and maintain benchmark test datasets to track, measure, and score AI response quality over time.
- CI/CD Integration: Integrate automated AI quality and regression checks directly into the continuous integration and deployment (CI/CD) pipelines.
- Bug Documentation & Escalation: Identify, isolate, and document AI-specific bugs, model drift, and system anomalies with clear reproduction steps for Machine Learning and Engineering teams.
Key Requirements
- Experience: 3+ years of experience in Software Quality Assurance, with a focus on Conversational AI, Chatbots, or LLM-powered applications.
- Testing Focus: Proven track record in testing natural language processing (NLP) models, prompt safety, and hallucination evaluation.
- Automation Skills: Hands-on experience with API testing (Postman, REST Assured, or Python) and UI automation tools (Selenium, Playwright, Appium, or Cypress).
- Technical Proficiency: Experience with CI/CD tools (GitHub Actions, GitLab CI, or Jenkins) and scripting languages (Python or JavaScript).
- Adversarial Knowledge: Familiarity with common LLM vulnerabilities, prompt injection techniques, and security testing principles.
- Communication & Detail: Strong analytical and problem-solving skills with the ability to communicate complex model behavior issues clearly to engineering teams.
Independent Contractor Perks
- Permanent work from home
- Immediate hiring
- Health Insurance Coverage for eligible locations
Note
Please click the Apply Now button to complete your application, including the assessment questions, technical check, and voice recording. Your hourly pay rate will be established based on your performance in the application process; submissions with all requirements fulfilled will receive priority review.
One application, multiple possibilities: When you apply, our team reviews your background against all current openings, not just this one! If a different role fits your skills, we'll get in touch! We also encourage you to keep exploring our job board and apply directly to any position that excites you.
Reminder
Apply directly to the link provided; you will be redirected to BruntWork’s Career Site. You must apply using the link to complete the initial requirements, which include pre-screening assessment questions, a technical check of your computer, and a voice recording. APPLICATIONS WITH COMPLETE REQUIREMENTS WILL BE PRIORITIZED.