Porter AMR: UX Research & AR Testing
Pioneering the UX of autonomous mobile robots for the warehouse — using AR to test robot communication and behaviour before development.
Autonomous mobile robots are entering real workplaces. But no UX playbook existed for how they should communicate with the humans working alongside them. Co-leading with a UX researcher and working with a senior robotics engineer, we ran a three-phase research programme to define how the Porter AMR should behave, signal, and communicate in a shared human workspace. The output: UX principles and interaction guidelines handed directly to the robot engineering team. This was the first AR UX test conducted on the company's B2B side.
- Role
- Co-lead on research — research planning and testing design
- Team
- Senior UX designer, senior UX researcher, senior robot engineer, UX grad
- Timeline
- August 2025 – January 2026
- Tool
- Meta Quest 3

The Problem
The robotics team was preparing to deploy Porter to client warehouse sites and needed interaction guidelines for some important decisions: how Porter should behave alongside human operatives through its lighting, sounds, signals, and movement, how it should ask for help when stuck, and how it should grab an operative's attention when it needs help or warn them to protect their safety.
And one more ambitious question: could giving Porter a personality help workers accept it?
Our research questions
- What lighting, audio, and visual signals help workers understand what Porter is doing?
- What communication builds trust and makes Porter's behaviour predictable?
- How should Porter signal that it needs help, and will workers respond to it?
- Does giving Porter a personality improve collaboration, or is it a distraction?

The Approach
As the first AR test on the company's B2B side, we planned three phases for this project, each narrowing the focus before committing to the next:
- 01Stakeholder workshops — two online groups mapped the territory and defined the six use cases that structured everything after: obstacle encounters, help requests, emergency stops, pass-bys, tight manoeuvres, direct approaches.
- 02Office testing — participants rated communication elements per use case, narrowing hundreds of combinations down to the strongest candidates.
- 03AR testing in the warehouse — the top candidates, tested at full scale in a realistic environment with a virtual Porter.
Before the workshops, we ran a supervised trial of Porter on the fulfilment centre floor and analysed the video footage, categorising each human-robot interaction by the intent signal or behavioural adjustment it called for. That, combined with a review of existing HRI approaches at Waymo, Uber, and Starship, gave the workshops a working set of communication elements to test rather than a blank page.
We chose AR for the final phase because a virtual robot let us swap lights, sounds, speeds, and stop distances between conditions in minutes: iterations that would have required enormous engineering effort on real hardware. It also let us safely test scenarios like close approaches and emergency stops that would carry genuine risk with a physical robot.

Phase 1 — Stakeholder Workshop
We gathered all stakeholders for an online workshop, split across two days with two groups of attendees each day. The workshop's goal was to explore and brainstorm the elements Porter would need to best communicate with warehouse operatives. We gave both groups a set of scenarios and asked them to rate how much influence they thought each element would have on those scenarios.
The workshop gave us a first set of elements to test later, along with ideas for how they might combine, and input from stakeholders across the business: product managers, warehouse managers, operatives, and engineers.

Phase 2 — Office Testing
We hypothesised that personality would build trust. The data said otherwise.
Participants were colleagues who had mostly never interacted with Porter, deliberately chosen to simulate first contact. For each of the six use cases, they rated combinations of lighting colour and frequency, audio, on-screen text, and personality expressions. A mock application showed exactly how lights would appear on the real robot, changing in real time as scenarios played out. Alongside it, 3D-printed models of Porter gave participants something physical to react to: a cheap, safe way to test how people believed they'd respond before we staged it for real in AR.


The personality finding was the most significant result. We expected attendees to rate personality — a face, character expressions — highly, believing it would make them more comfortable and more willing to help when Porter needed assistance. They didn't. What they consistently preferred was functional clarity: text on screen combined with audio. Not a character. Not a face. Just: what is this robot doing, and what do I need to know?
This single finding redirected the AR phase entirely, away from personality expression and towards communication hierarchy and signal clarity.

Phase 3 — AR Testing in the Warehouse
We built a virtual Porter that participants could see and react to in the real world simply by putting on a headset, working alongside it as they would the physical robot.
We ran the test at the company's internal testing centre, a facility replicating a real warehouse environment. Participants were recruited with prior AR/VR experience, so their reactions reflected Porter's behaviour rather than unfamiliarity with a headset.
Each participant completed three physical tasks: moving objects on a table, moving objects between tables, and sorting product orders — tasks chosen to occupy their attention the way real work does. Porter would then appear in the distance or drive towards them, lights and audio active, passing by or stopping at defined distances while they kept working. After each session, we interviewed participants and let them wear the headset again to see more element changes, so we could get their opinion directly.




The most valuable feedback method was also the simplest: participants could point directly at the virtual robot.
Rather than describing preferences in words, they gestured: “this light here, brighter, faster.” Pointing at something only they could see, in three-dimensional space, gave us precision that rating scales couldn't.
We captured video, screen recordings of Porter's appearance against human reaction timing, and real-time observation of whether participants noticed Porter at all: whether they looked up, changed gaze, or stayed absorbed in their task.
Key Findings
Trust comes from predictability, not personality.
Workers consistently preferred on-screen text plus audio over any character or expression. Clear state communication beat likability every time.
Audio needs a hierarchy, not a single signal.
Workers absorbed in a task don't always notice a passing robot. Volume and urgency must scale with the awareness required: a routine pass-by and an emergency stop need entirely different audio.
Lighting is the ambient channel; text and audio own the attention moments.
Distinct lighting patterns let workers read Porter's state without stopping. Combined with audio and text for higher-stakes scenarios, clarity improved significantly.
Stop distance matters as much as lights or sounds.
Too close felt threatening; too far looked uncertain. The right distance varies by use case and must be defined per scenario, not globally.
AR testing opened up opportunity across the company.
Many participants said the testing felt interactive and realistic enough to show them what a project would actually feel like in the real world, saving time and effort building something that wouldn't have worked. Almost everyone said they wanted this kind of testing on their own projects.
Output
The programme produced a detailed guidelines document for continued Porter development: recommended lighting colours and frequencies per use case, an audio urgency hierarchy, on-screen text recommendations, stop distances per scenario, a communication hierarchy across all three channels, and a “what not to do” section for elements that consistently tested poorly.
Every recommendation was evaluated against two criteria: worker safety and operational efficiency. Ultimately, the goal is to build a robot that workers understand and find easy to work alongside in the same environment.
The guidelines aren't just for Porter. They're the foundation for how the company approaches all future AMR UX: as the fleet grows, new work builds on this research rather than starting from scratch.
Reflection
The three-phase structure was right. Each phase earned the next investment, and AR delivered: the quality of feedback and environmental realism made it significantly more reliable than office testing alone.
Two things I'd change
More time in the warehouse.
The AR phase ran in a single day, and given how much the environment changed participant responses, its value was disproportionately high. I'd push hard for more use cases, more participants, more varied scenarios.
Capture data earlier than feels necessary.
A planned headset reaction-time measurement — precise, objective timing of how quickly participants noticed Porter — was cut by company restructuring before we could run it. It would have significantly strengthened the quantitative findings. Research timelines are vulnerable to organisational change in ways product timelines aren't.
The natural next phase: testing with real Porter hardware in a live deployment, and a longitudinal study of how trust develops as workers spend weeks alongside the robot. These are questions this research couldn't reach.