QA Tester, Conversational & Vision AI (Internship)
Job type: Full-time internship (paid)
Location: Kuala Lumpur, Malaysia (hybrid)
Please note: This role requires a high level of written and spoken English — roughly IELTS 7.0 / CEFR C1 / MUET Band 4.5 or above. All testing, conversation, annotation, and reporting is conducted in English.
About the role
We are developing conversational and vision AI products, and need a QA tester to evaluate both: testing whether our conversational AI sustains context and recalls information accurately across sessions, and whether our vision model's analysis of video footage matches what the footage actually shows — noting the correct detail and timestamp wherever it doesn't. The role also includes data labelling. Both halves reward the same instinct: noticing when something is slightly off and being unable to let it go. The ideal candidate is a sharp communicator and storyteller with a good eye and memory, not an engineer.
This role is for you if you're:
- A psychology, cognitive science, or human-computer interaction student — someone who thinks about how people are understood, remembered, and perceived
- A linguistics, communications, or humanities student drawn to language and conversation
- A creative or storytelling type — writing, film, theatre, or the arts — who is good at drawing out a conversation and building a narrative, and who watches footage closely
- Anyone fascinated by memory, human relationships, perception, and what makes a conversation feel real
You do not need to be an engineer or have a technical background. You do need excellent English.
Responsibilities
(1) Conversational AI testing
- Conduct rich, sustained test conversations with the AI in English, drawing on detailed narratives and observations to evaluate conversational quality
- Assess whether the AI maintains context within a conversation and recalls information accurately across sessions
- Track precisely what was said and when, and identify inconsistencies or errors in the AI's recall over time
- Evaluate the memory-review and consent experience for clarity and accuracy
(2) Vision model QA and data labelling
- Watch assigned video footage attentively and in full, often more than once
- Review the analysis summaries produced by the vision model against what actually occurs in the footage
- Identify discrepancies — missed events, misidentified objects, people or actions, incorrect sequencing, hallucinated detail, wrong counts, or errors in timing
- Record the correct information along with the precise timestamp of each discrepancy, following our annotation format and conventions
- Label and annotate video and image data accurately and consistently, applying labelling guidelines as they evolve
- Flag ambiguous or edge cases rather than guessing, and raise questions where the guidelines don't cleanly cover what you're seeing
Must-have qualifications
- Excellent written and spoken English, approximately IELTS 7.0 / CEFR C1 / MUET Band 4.5 or above — you will be conversing, annotating, and writing reports in English all day, and subtle errors in the product can only be caught by someone with a strong feel for the language
- Ability to conduct rich, sustained conversations and storytelling, bringing detailed narratives and observations into the interaction
- Strong ability to accurately remember and track what was said and when, in order to detect inconsistencies in the AI's recall
- Careful visual attention — able to watch video closely for extended periods and notice detail that is easy to skim past
- Precision and consistency in record-keeping, including accurate timestamping and adherence to a labelling convention
- Comfort with repetitive, detail-heavy review work, and the patience to stay accurate on the fortieth clip as much as the first
- Willingness to share experiences and observations during testing, on a voluntary basis, to enable realistic evaluation
Data and privacy
Any personal information you choose to share during conversational testing is used solely to evaluate the product. Participation is voluntary, and we will explain how your data is stored, who can access it, and how you can request its deletion.
Video footage reviewed as part of this role may contain identifiable people and other sensitive material. All footage and analysis output are confidential, must be reviewed only on approved systems, and may not be copied, shared, or discussed outside the team.