Engineering

Human-AI QA Engineer (Mid level)

You will test AI systems where the correct output is not a fixed string: writing evaluation suites for model behaviour, building regression harnesses that catch drift between model versions, and finding the failure cases that only show up under real user pressure. You sign off on what ships.

5-8 YearsHyderabad, TG, INHybrid/RemoteFull Time

The work

Conscipact builds its own AI products and builds AI systems for enterprise clients. Both are judged the same way: whether the person on the other side of the system is better off using it. That makes QA harder than pass/fail. A response can be fluent, fast, well-formed and still wrong for the person who asked. Your job is to find that gap before a user does, and to write it down in a way an engineer can act on.

What you'll do

Own the test strategy for AI-backed features end to end, from unit and integration coverage of the surrounding application to evaluation of the model layer itself. Build and maintain evaluation sets with known expected behaviour, including adversarial and edge-case inputs, and run them on every model or prompt change so regressions are caught before release, not after. Test non-deterministic output. Define what "correct" means when there are many acceptable answers, and write assertions that hold up against that definition rather than against one lucky sample. Break things deliberately: prompt injection, malformed and hostile inputs, long-context degradation, tool-call failures, timeouts, partial responses. Check the explainability requirement in practice. When a system produces an output a user would want to question, verify the reasoning or provenance it exposes is actually there and actually accurate, not a plausible-sounding after-the-fact narrative. Verify that decisions the system routes to a human are in fact routed to a human, and that the handoff carries enough context for that person to decide rather than rubber-stamp. Write defect reports with reproduction steps, frequency and severity. Argue for the ones that matter when a release is under time pressure.

What we're looking for

Five to eight years in QA or test engineering, with real ownership of quality for a production system, not just execution of someone else's test plan. Strong automation skills in Python or JavaScript. You can write and maintain a test harness, not only use one. Experience testing APIs, and comfort with CI pipelines to the point where you can add and debug your own test stages. Some direct experience testing or evaluating LLM or ML-backed features. This can be at work or self-driven, but you should be able to describe how you decided what "good" looked like and how you measured it. Willingness to say a release is not ready and to hold that position with evidence.

How we work

Two frameworks run through every deliverable. Awareness-First AI is the standard every system we ship is held to, in this order: context before capability; wellbeing is the metric, meaning not engagement, not retention, not task volume, not tokens served; judgment stays human, so consequential decisions are routed to a person rather than automated away; explainable under pressure, meaning the system can account for a specific output to a specific user on the day it goes wrong, not in a whitepaper. Conscipact HAI is how we work internally. AI carries volume, humans carry judgment. Every output has a named human accountable for it. For your work, that means you can use models to generate test cases, fuzz inputs and triage logs, and you remain the person answerable for what the suite did and did not catch. We are not an AI consultancy with a values page, an ethics advisory firm, or a wellness brand. We ship production systems and these principles are testable properties of them, which is why this role exists.

What we'll ask you

No surprises — these are the questions in the application, so you can think about them before you start.

  • Describe a feature you tested where the output was not deterministic. How did you define correct, and what did your assertions actually check?
  • Tell us about a defect that reached production on your watch. What was it, why did your tests miss it, and what changed afterwards?
  • Have you ever blocked or delayed a release over a quality issue when there was pressure to ship? What was the issue and what happened?
  • Which language do you write your test automation in, and roughly how much of your recent test suite did you write yourself?

Apply for Human-AI QA Engineer (Mid level)

Four short steps — about ten minutes. You'll need a résumé (PDF, DOC or DOCX, up to 4MB). A person reads every application, and you’ll get a reference number so you can check where it stands.

Start your application

See the other open roles, or leave your details at the Open Door if this isn’t quite it — we look there first when we write the next one.