Member of Technical Staff

 

Description:

We’re building Buddy: a personal intelligence that learns one specific person deeply enough to increasingly think, decide and act the way they would at their best, while remaining bounded by their own values.

Not a chatbot.

Not a copilot waiting to be prompted.

The end state is something closer to a persistent extension of the person themselves.

I’m looking for someone who wants to work on the intelligence itself.

 

The Problem

Current AI systems are remarkably capable. But what we are trying to build requires more than putting a frontier model behind tools.

Buddy has to live alongside one person for years, learn from what happens, and become more capable through that experience.

That creates problems I do not think the field has solved.

How should an intelligence learn continuously from years of decisions, actions, corrections and outcomes without becoming unstable or forgetting what it already knows?

When should an experience become memory, a preference, a learned skill, a policy change, a model update, or nothing at all?

How do you teach judgment from longitudinal real-world outcomes rather than only static preference data?

How does an agent learn whether an action worked for the right reason, assign credit when consequences arrive much later, and improve how it reasons and acts the next time?

We will use existing methods wherever they are good enough.

Open-weight models. Fine-tuning. Post-training. Reinforcement learning. Preference optimisation. Verifiers. Distillation. Test-time compute. Continual learning. Agent learning.

But I do not assume today’s methods are sufficient.

Where they are not, I want us to find out why and build something better.

That is the job.

 

What you would work on

Your primary technical domain is Models & Learning.

That includes model training and post-training, reinforcement and preference learning, continual and online learning, reasoning, agent learning, model adaptation, evaluation, reward models and verifiers, synthetic environments, inference-time learning, specialist models, and new learning methods where existing approaches are insufficient.

This is deliberately not a fixed research roadmap.

If the most important thing we are working on eighteen months from now does not have a name today, that is a good outcome.

You will also build.

An idea here has to survive the path from hypothesis to experiment to code to training run to evaluation to a working system.

I am not looking for a structure where one person researches and somebody else turns it into reality.

 

Who I am looking for

I care far less about your current title, age, or number of years worked than I do about evidence of unusual technical ability.

You may come from a research lab. You may have a PhD. You may not.

You may be a research engineer whose strongest work never became a paper.

You may have built important open-source systems or models.

You may simply have spent the last few years going unusually deep on hard problems.

What matters is that you can go beyond implementing what already exists.

You understand modern machine learning deeply enough to question its assumptions.

You can read research, reproduce it, break it, modify it and design the experiment that tells you whether your alternative is actually better.

You have research taste: you can distinguish something technically interesting from something that actually matters.

You can reason from first principles when there is no established playbook.

You are as comfortable opening a codebase as opening a paper.

You do not need a fully specified ticket before you can start thinking.

And you are willing to tell me when an idea is technically wrong, including mine, and do the work required to prove it.

Organization WYLE
Industry Other Jobs Jobs
Occupational Category Member of Technical Staff
Job Location Dubai,UAE
Shift Type Morning
Job Type Full Time
Gender No Preference
Career Level Intermediate
Experience 2 Years
Posted at 2026-09-21 6:20 pm
Expires on 2026-12-20