# Abdul Rashid — Blog > Articles by Abdul Rashid on browser clouds, AI-agent evaluation harnesses, faster mobile proofs, and Playfield’s spatial computing experiments. ## Pages - [Blog](https://abdulrashid.dev/blog/index.md): Articles by Abdul Rashid on browser clouds, AI-agent evaluation harnesses, faster mobile proofs, and Playfield’s spatial computing experiments. - [Can you leave your agent alone for ten hours?](https://abdulrashid.dev/blog/can-you-leave-your-agent-alone-for-ten-hours/index.md): How repeatable evaluation harnesses help AI agents work independently, guide experiments, and deliver reliable results over long tasks. - [Why I’m building Playfield](https://abdulrashid.dev/blog/why-im-building-playfield/index.md): Why Abdul Rashid is building Playfield: hand-controlled games with an iPhone and projector, inspired by EyeToy, Party Fowl, and Folk Computer. ## Full content Complete HTML content converted to Markdown. Images retain their URLs, alt text, and captions. --- Source: https://abdulrashid.dev/blog/ Markdown: https://abdulrashid.dev/blog/index.md # Blog Notes from building infrastructure, giving AI agents more autonomy, and exploring spatial computing. ## [Why I’m building Playfield](https://abdulrashid.dev/blog/why-im-building-playfield/) From EyeToy with friends to games on a wall. The story behind Playfield, inspired by Party Fowl, Folk Computer, and the joy of playing with your body. ## [Can you leave your agent alone for ten hours?](https://abdulrashid.dev/blog/can-you-leave-your-agent-alone-for-ten-hours/) How evaluation harnesses let AI agents work with more autonomy, with case studies from XPath-Go, Popcorn, and our provider-creation agent. ## [Why We Built Popcorn: An Attestable Browser Cloud for Reclaim](https://blog.reclaimprotocol.org/posts/why-we-built-popcorn) Building our own browser cloud: confidential computing, a regional fleet, and a browser experience that feels natural on a phone. ## [Turbocharged Zero-Knowledge Proofs for Mobile](https://blog.reclaimprotocol.org/posts/gnark-migration) Moving mobile proof generation from a WebView to native Go and gnark, cutting average proof time from about 40 seconds to 4–5 seconds. --- Source: https://abdulrashid.dev/blog/can-you-leave-your-agent-alone-for-ten-hours/ Markdown: https://abdulrashid.dev/blog/can-you-leave-your-agent-alone-for-ten-hours/index.md # Can you leave your agent alone for ten hours? By Abdul Rashid · 6 min read ![A hand-drawn rider guiding a horse with a harness]() Getting a PR from a reported issue is already something AI models can handle pretty well. But give them a long horizon task and they start to struggle, lose direction, and take shortcuts. If we have to keep looking over the agent’s shoulder to check its work, we are still the bottleneck. If we want to give agents more autonomy, we need to create a harness. By harness, we mean an evaluation loop around the task, beyond tools like Codex and Claude Code. It should emulate how we would verify the work and guide the agent: try it, find what is wrong, and give it enough feedback to know what to do next. That lets it keep moving without waiting for us after every change. Some case studies from what we have been working on. ## XPath-Go Moving to the TEE stack meant we needed an XPath library in Go. The agent got parts of it working and the Go tests were passing, but the results differed from our JS version. A simple comparison test gave us a way forward: give both versions the same input and compare what comes back. This grew into a suite comparing our Go library against jsdom, checking nodes, text, namespaces, and source locations. That made a big difference. The agent could see exactly where the two implementations disagreed and keep working through those differences. I could give it a goal, go off for around ten hours, and come back to a much better implementation. The recorded run passed **888 out of 888 compatibility cases**. There are still things outside that test coverage, but we now had a concrete way to measure progress without checking each fix ourselves. [Results](https://github.com/reclaimprotocol/xpath-go/blob/main/docs/COMPATIBILITY.md) ## Popcorn The goal with Popcorn was to build the best mobile webview. That meant testing how it performs across a wide variety of devices, operating systems, and websites. Even figuring out what to test is hard: every website has different interactions, and those can behave differently depending on the screen size, OS, or keyboard. Turning all of that into a fixed set of automated tests is a lot of work. So we built a flow where the agent explores interactions, creates test cases and repeatable test pages, then compares the same interaction directly in the device browser and through Popcorn. We wanted the testing to be close to how we would do it ourselves. So it uses native taps and gestures at coordinates and captures the actual screen. If an input is hidden behind the keyboard, we need to see that. Knowing that the browser thinks the input is focused does not tell us whether someone can use it. We test across iOS, Android, and multiple keyboards. We also started putting parts of the checking into code. Pixel diffs let us repeat visual checks with defined thresholds, without spending AI tokens looking at every screenshot on every run. The agent can spend its time finding new cases and investigating failures. [Harness](https://github.com/reclaimprotocol/popcorn-oss/blob/main/images/minimal-vnc-desktop/mobile-harness/README.md) ## An AI agent improving an AI agent (AIception?) For our provider-creation agent, the goal was to have AI improve it, try new providers, and compare different models and approaches. The harness was an evaluation setup with 22 education portals under our control. With known accounts, test data, and expected results, we could run each variation through the same tasks and see what worked. Creating a provider was only part of the test. It had to return the logged-in user's complete name, not just their first name, and the proof could not leak their username or password. The generated injection also had to detect login correctly, keep the verification overlay hidden before login, and show it afterwards. These were the things we would check ourselves, now built into the evaluation. The agent still found ways to bend the rules. In one experiment, it added a provider fixer during evaluation. The task was to create a provider and check whether it could replay successfully. Fixing it along the way made the result look better, but did not tell us whether the original provider worked. So the repair step was removed: the harness pins the exact version created and replays it in a fresh session with AI disabled. Both creation and replay have to pass. [Playback contract](https://github.com/reclaimprotocol/education-portal-evals-infra/blob/main/docs/provider-playback-validation.md) More autonomy still needed some human nudges. Sometimes the agent would keep trying variations of the same approach when it needed to look somewhere else. Sharing what we already knew about the system and giving it access to sister repos helped it do that. Some failures led to Portal filtering valid JSON responses; others involved the TEE getting an empty response body. With that broader context, it could investigate the actual problem across repos instead of trying to fix everything inside the agent repo. ## More trust, more autonomy Agents will keep getting better, but if we are still checking everything they do because we don’t trust them, we will miss out on a lot of those gains. Our time is better spent improving the checks, sharing context, and giving a new direction when needed. Letting the harness handle repeated testing builds the trust to give agents more autonomy. It also lets us run more agents in parallel without every result queuing up for us to check. --- Source: https://abdulrashid.dev/blog/why-im-building-playfield/ Markdown: https://abdulrashid.dev/blog/why-im-building-playfield/index.md # Why I’m building Playfield A wall, an iPhone, and the joy of playing with your body. By Abdul Rashid · 5 min read Some of my favourite childhood memories are of playing EyeToy with friends. A camera plugged into the PlayStation, a game on the TV, and suddenly our bodies were the controllers. There’s a particular feeling you get from that kind of play. You reach for something and the game responds. You move a little too much, look a little ridiculous, and your friends laugh. The fun spills out of the screen and into the room. Even watching someone else play becomes part of it. [EyeToy](https://blog.playstation.com/2010/11/03/eyetoy-innovation-and-beyond/) made that feel possible with a little camera. What stayed with me was how much fun it was to move around and play together. ## That feeling came back Years later, playing [Party Fowl](https://www.nex.inc/partyfowl) brought it back. It uses a device’s camera to turn body movements into controls for wonderfully silly games. There was that familiar feeling again: moving, reacting, and laughing at what the game was asking us to do. It reminded me how much I like games that give you a reason to get up. The movement is part of the pleasure. So is being in the same room as the people you’re playing with. That’s the feeling behind [Playfield](https://projectplayfield.com/), an experiment I’m building with an iPhone, a projector, and games you play with your hands. The game goes on the wall. Your hands do the rest. ![A player reaching up to move a projected paddle in Playfield’s Meteor Keeper game]() Meteor Keeper. One wall, a small city, and two very busy hands. ## Let the room join in Spatial computing has interested me for a long time, especially mixed reality, augmented reality, and tangible interfaces. I like the possibility of reaching into a digital experience through the space and objects around us. [Folk Computer](https://folk.computer/) was a big inspiration. It uses cameras and projectors to bring computation onto physical surfaces and objects. Paper can become an interface. Moving something on a table can change what a program does. The computer gets a place in the room, alongside the people using it. What drew me in was how much room that leaves for experimentation. An interface can be something you touch, rearrange, or share with another person. Seeing Folk made me want to explore that direction myself. Playfield starts with a small piece of it: a wall that responds to your hands. Something you can walk up to, understand by trying, and play together. ## A smaller setup Interactive walls and floors have existed for years. Systems like [LUMOplay](https://www.lumoplay.com/) already turn projected surfaces into places to play. Their [recommended setups](https://www.lumoplay.com/hardware) combine a computer, a projector, and a dedicated 3D camera. That made me curious about what a phone could take over. A modern iPhone brings a camera, the compute to run hand-tracking models, and, on supported models, LiDAR depth sensing into one device. AI vision models can locate hands in the camera image; depth can help with interactions that need to know how close a hand is to a surface. For this experiment, the phone handles tracking and runs the games. The projector makes them big enough to play on a wall. LiDAR is used in the wall-touch mode; the other games use camera-based hand tracking. There’s still a projector to connect and a phone to mount. But being able to try this without a separate computer and tracking camera makes it much easier to keep experimenting. ![An iPhone mounted beneath a projector on the Playfield rig]() The current rig: a projector, an iPhone, and a mount to keep them still. ## A wall with a few new rules The phone first looks for markers projected at the corners of the play area. That lets it map what the camera sees to where things appear on the wall. Once calibration is ready, the markers disappear and the game takes over. In Meteor Keeper, your hands steer paddles to bounce meteors away from a city. In Counter Crew, you grab tomatoes, chop them with your hand, cook imaginary soup, and serve it. There’s a drawing mode for leaving trails across the projection, and a physics playground for grabbing and dropping a ball. The games are deliberately small. They’re ways to find out which interactions feel good: reaching, grabbing, opening your hand to let go. When the hands leave view, the games can pause. Even an imaginary kitchen should let you take a break. ![A player using an open hand to interact with a soup pot in Playfield’s projected Counter Crew kitchen]() Counter Crew. Chop, drop, and cook. The soup is imaginary; the arm movements are real. ## What I want to build toward Playfield is still an experiment. A lot of the work is making the tracking, calibration, and gestures feel natural enough that you can pay attention to playing. A hand movement should do what you expect, without making you think about the camera watching it. The goal is a setup that’s easy to bring out, with games that invite people to join in. There’s plenty to explore beyond these first few games, but this feels like a good place to start. EyeToy gave me some of my favourite memories with friends. Party Fowl reminded me how good that kind of play can feel. Folk opened up more possibilities for what the space around us could do. Now there’s a little city on my wall, and it needs saving. [See Playfield in action →](https://projectplayfield.com/)