Senior AI Quality Engineer
Job Description Roles & Responsibilities Test automation across the stack Backend: API and contract testing, service-level and integration coverage, data setup that doesn't rot, and test design that survives a schema change. Frontend: web E2E and component-level coverage with Playwright, visual and RTL regression, and suites fast enough to gate a merge rather than a nightly. Mobile: native and cross-platform coverage with Appium or Maestro, device-farm strategy, offline and sync behaviour, and the payment-peripheral paths that only break on real hardware. The connective tissue: shared fixtures, environment and test-data management, parallelisation, and CI pipelines where a red build means something. performance and load testing Experience. The AI layer on top of it Test generation from specs, code, and production traffic with the maintenance story solved, not just the first draft. Failure triage that classifies a red build before a human opens it: real bug, flake, environment, or test rot. Self-healing locators and suite health tooling flake detection, quarantine, coverage-gap analysis. Evaluation infrastructure for AI features across our products: datasets, scoring, and regression detection when a prompt or model changes. Evaluation for our market specifically Arabic and English behaviour, RTL interfaces, and region-specific POS, tax, and payment rules. Correctness here is rarely a string match. Agentic AI and orchestration Agentic AI that does real work in our pipelines: reads a diff, runs the relevant suite, reproduces a failure, proposes a fix, opens the PR. Agents that own a quality workflow end to end exploratory testing against a running build, coverage-gap hunting, release-risk assessment and know when to escalate to a human. Orchestration that holds up under load multi-step planning, tool use, retries, state and memory across steps, sandboxed execution, multi-agent handoffs, and clean boundaries between agentic and deterministic steps. Integration with the stack we already have (CI, Jira, observability, MCP-style tool interfaces) rather than a parallel system beside it. The judgement to know when a plain pipeline beats an agent, and to say so. The technical ground You should be current on how this work is actually done today, and able to argue about it rather than recite it: Test automation: framework design and layering, the test pyramid and where it stops being useful, flake economics, parallel execution, mobile and cross-browser realities, CI/CD gating, Framework: Playwright, Appium, Maestro Agentic AI: orchestration and tool use, multi-step planning, memory and state, sandboxed execution, multi-agent patterns, MCP and similar tool-integration standards, and the cost of each. Context engineering: retrieval strategy, chunking, reranking, caching, and managing long-context behaviour including where it degrades. Evaluation: offline and online evals, LLM-as-judge and its failure modes, human-in-the-loop review, statistical significance on small samples, regression gates in CI. Reliability: structured output, guardrails, fallback and retry design, and handling non-determinism in systems that must not flap. Operations: tracing and observability for LLM systems, prompt and version management, latency and cost budgeting, model routing, and when fine-tuning or distillation beats a better prompt. Desired Candidate Profile An engineer who ships production software, with recent hands-on work on LLM-backed systems that real users depend on. Strong Python; comfortable in at least one of .NET, Java, or TypeScript. Tested, maintained code not notebooks. Real automation depth across more than one surface. You've owned a suite that gates releases on backend and on a UI web or mobile and you can explain how you kept it green without deleting the hard tests. Real experience building evaluation systems. You can explain how you knew your system was getting better, with numbers. Practical depth with the modern LLM toolkit prompting, structured output, tool use, retrieval, agentic AI orchestration and a clear sense of the trade-offs. Credible testing fundamentals. You don't need a QA title, but test design, automation frameworks, and CI/CD shouldn't be new to you. A bias toward adoption. You measure your work by what other engineers use, not by what you demoed. Company Industry IT - Software Services Department / Functional Area IT Software Keywords Senior AI Quality Engineer Get real-time job updates only on our App
Ready to apply?
You are viewing this role on JobSphere AI. Applications are completed on the original employer / source website.
Apply NowOpens the employer's site in a new tab
- CompanyFoodics
- LocationCairo, Egypt
- CategoryFullStack
- SourceNaukrigulf
- Listedjust now
Related FullStack jobs
Phlebotomist
Perform venipuncture and capillary blood collection in patients homes in accordance with approved policies, procedures, and infection prevention standards…
Full Stack Engineer
Python Development: Deep knowledge and development skills in Python and its problem required problem solving capabilities. Front End Development: Collaborate…
Senior Business Analyst Analytics, Automation & AI
About Amazon Now MENA Amazon Now operates across the UAE, Kingdom of Saudi Arabia, and Egypt, delivering fast and customer-focused commerce solutions. The…
Developer
Design, develop, and maintain software applications according to specifications Collaborate with product managers, designers, and other stakeholders to define…