{"version":1,"url":"https://www.piyushvyas.com","profile":{"name":"Piyush Vyas","bio":"Piyush Vyas builds AI products and leads engineering teams. Explore Agentic Intelligence, work at AWS and Numerade, research, writing and contact information.","email":"work@piyushvyas.com","url":"https://www.piyushvyas.com/","profiles":[{"label":"GitHub","url":"https://github.com/pyyush"}]},"pages":[{"id":"about","title":"About Piyush","kind":"information","url":"https://www.piyushvyas.com/about/","summary":"Meet Piyush Vyas, founder of Agentic Intelligence. Background in AI products, engineering and delivery, with contact details and professional links.","text":"Piyush Vyas AI products, from the first idea to the team that builds them. work Building AI products & teams now Founder · Agentic Intelligence email work@piyushvyas.com Copy email address experience Résumé & awards work Products I’ve worked on writing What I’m building & learning research Papers & awards GitHub pyyush website piyushvyas.com about this conversation A portfolio told through messages. Not a live chat with Piyush. privacy Email drafts stay on your device"},{"id":"work","title":"Selected work","kind":"information","url":"https://www.piyushvyas.com/work/","summary":"Selected AI product and engineering work by Piyush Vyas: Agentic Intelligence, AWS delivery, Ace AI Tutor, AI video, moderation and speech recognition.","text":"Work A few things I’ve helped build Agentic Intelligence Jun 2026–Present AI delivery at AWS Sep 2024 – May 2026 Ace AI Tutor Apr 2023–May 2024 AI video platform Apr 2023–May 2024 Voice & text moderation Oct 2021–Apr 2023 Speech recognition Aug 2021–Oct 2021"},{"id":"writing","title":"Writing","kind":"information","url":"https://www.piyushvyas.com/writing/","summary":"Technical writing by Piyush Vyas on browser agents, harness engineering, deterministic tooling, semantic element identity and applied AI.","text":"writing Browser agents & applied AI How I Would Design a Tiny Parking Agent for Visitors Design concept · 4 min Harness Engineering Agent reliability · 3 min Deterministic Tooling for Agent Skills Public tooling · 6 min Booking Open Play Before It’s Gone Browser agents · 5 min Browser Agent Protocol From the archive · November 9, 2025 Universal Semantic Element ID From the archive · March 12, 2026 I Replayed a Browser-Use Agent Session Deterministically From the archive · March 26, 2026"},{"id":"resume","title":"Piyush Vyas — résumé","kind":"information","url":"https://www.piyushvyas.com/resume/","summary":"Piyush Vyas’s experience résumé: roles, responsibilities, outcomes and research awards across Agentic Intelligence, AWS, Numerade, Spectrum Labs and Comcast.","text":"Piyush Vyas Experience résumé work@piyushvyas.com · GitHub Agentic Intelligence LLC Founder Jun 2026–Present Founded Agentic Intelligence to design and deploy AI agents for customer-facing and operational workflows. Build browser-use, computer-use, and conversational agents across voice, messaging, email, and web. Amazon Web Services (AWS) AI Delivery & Engagement Lead Sep 2024 – May 2026 Led 40+ AI/ML and generative AI engagements across AMER and EMEA, serving startups through Fortune 500 companies in financial services, software & internet, healthcare & life sciences, and automotive. Led multidisciplinary teams of consultants, partners, account teams, and customer engineers through discovery, architecture, prototyping, evaluation, and production readiness, managing scope, milestones, delivery risk, and executive communication. Led 25+ executive workshops with C-suite, business, and engineering leaders to prioritize use cases, define technical and business KPIs, and turn broad AI goals into practical delivery roadmaps. Designed and prototyped agentic systems, including browser-use, computer-use, and conversational agents, alongside RAG and computer vision; turned recurring patterns into reusable architectures and accelerators for AWS teams and partners. When customers moved to production, partnered with AWS Professional Services and customer production teams to support deployment, then led adoption, enablement, and stakeholder alignment through rollout. Numerade Labs, Inc. AI Product & Engineering Lead Apr 2023–May 2024 Partnered directly with the CTO and Head of Product to shape the roadmap and technical direction for Ace AI Tutor and Numerade’s AI video platform, turning student behavior, educator feedback, and business goals into product priorities. Built and led a six-person ML team—4 engineers and 2 interns—setting technical direction, managing delivery, and promoting one engineer to senior within 12 months. Architected Ace as an agentic tutoring system using GPT-4, tool and function calling, image generation, and sandboxed Python code execution; built an automated grader and feedback loop to improve prompts, tools, and answers. Added visual problem-solving to Ace, driving 45% higher student engagement, more follow-up questions, and stronger retention. Built a multimodal video-generation pipeline using image generation, ElevenLabs text-to-speech, and HeyGen avatars, producing videos near-instantly at roughly one-fifth the cost of educator-created content; rollout across 50+ academic partners reached 90%+ adoption and contributed to growth in paid users. Spectrum Labs, Inc. Senior Machine Learning Scientist Oct 2021–Apr 2023 Built a voice toxicity detection product from the ground up—created a proprietary labeled audio dataset and trained voice-native models to identify abuse, bullying, hate, and other harmful behavior in live conversations. Worked with the CTO and engineering team to turn the research into a product sold to multiple customers, then expanded it from English to 12+ languages using multilingual speech models, including XLS-R 53. Took ownership of text moderation, moving the stack from classical classifiers to transformer models—training custom models from scratch, continuing pretraining on proprietary data, and deploying them through ONNX. Grew into the technical lead across voice and text moderation, working with the CTO and VP of Engineering on model strategy, evaluation, production integration, and the handoff from research to engineering. Joined customer and sales calls as the AI lead, translating trust and safety requirements into product capabilities and helping the team win and expand customer deployments. Comcast Applied AI Research Intern (NLP) Aug 2021–Oct 2021 Built a wav2vec 2.0–based speech recognition system for Comcast voice products, achieving a 30% reduction in word error rate for experiences serving millions of users. Co-authored research on temporal early exiting for low-latency streaming speech-command recognition, published at IEEE ICASSP 2022. Awards Recognition for coauthored research 🏆 Best Student Paper Award INTERSPEECH 2021 Optimally Encoding Inductive Biases into the Transformer Improves End-to-End Speech Translation Coauthored with Anastasia Kuznetsova and Donald S. Williamson. Read the paper 🏆 Outstanding Student Paper Award IEEE ICASSP 2021 An End-to-End Non-Intrusive Model for Subjective and Objective Real-World Speech Assessment Using a Multi-Task Framework Coauthored with Zhuohuang Zhang, Xuan Dong, and Donald S. Williamson. Read the paper Download experience (.txt) Research awards include linked attribution."},{"id":"privacy","title":"About this website","kind":"information","url":"https://www.piyushvyas.com/privacy/","summary":"How this portfolio works: local email drafts, private reactions, system fonts and no advertising or analytics scripts.","text":"not a live messenger The conversation presents Piyush’s career. It does not connect to Apple Messages or claim Piyush is online. contact Write an introduction to open a draft in your email app. You review and send it there. This website does not send, upload, or store your text. local interactions Reactions are just for you and reset when you reload. Nothing is sent to Piyush. The portfolio has no analytics or advertising scripts. typeface All interface text uses your device’s installed system fonts. No external font request is made. design reference Messages interaction design adapted from the sasi.codes archive supplied by Piyush."},{"id":"agentic","title":"Agents that work where people do","kind":"work","url":"https://www.piyushvyas.com/work/agentic/","summary":"I started Agentic Intelligence to build agents that can do useful work in the tools people already use: browsers, computers, voice, and messaging.","text":"Agents that work where people do Agentic Intelligence LLC I started Agentic Intelligence to build agents that can do useful work in the tools people already use: browsers, computers, voice, and messaging. Founder Jun 2026–Present What I worked on I’m building browser-use, computer-use, and conversational agents, across voice, messaging, email, and the web. More work"},{"id":"aws","title":"AI projects at AWS","kind":"work","url":"https://www.piyushvyas.com/work/aws/","summary":"I worked with customers across the Americas and EMEA to pick useful AI projects, build prototypes, and help their teams take them into production.","text":"AI projects at AWS Amazon Web Services (AWS) I worked with customers across the Americas and EMEA to pick useful AI projects, build prototypes, and help their teams take them into production. AI Delivery & Engagement Lead Sep 2024 – May 2026 40+ AI / ML engagements 25+ Executive workshops AMER · EMEA Delivery scope What I worked on I led 40+ AI/ML engagements and 25+ workshops with executives and their teams. We worked out which problems were worth solving, how to measure progress, and what to build first. I stayed hands-on with browser and computer-use agents, RAG, conversational AI, and computer vision. For production deployments, I partnered with AWS Professional Services and customer engineers. The 40+ figure includes workshops and prototypes—not just production launches. More work"},{"id":"ace","title":"Ace AI Tutor","kind":"work","url":"https://www.piyushvyas.com/work/ace/","summary":"We built Ace to help students work through a problem, not just see the final answer.","text":"Ace AI Tutor Numerade Labs, Inc. We built Ace to help students work through a problem, not just see the final answer. AI Product & Engineering Lead Apr 2023–May 2024 45% Higher engagement 6 ML team members GPT-4 Built with What I worked on I worked with the CTO and Head of Product on what Ace should do next, and led four engineers and two interns. We used GPT-4 with tool calling, image generation, and Python execution. I built a grader and feedback loop to improve its answers. Adding visual problem-solving increased student engagement by 45%. Tools used OpenAI Python More work"},{"id":"video","title":"AI video at Numerade","kind":"work","url":"https://www.piyushvyas.com/work/video/","summary":"At Numerade, we also built a way to make educational videos with AI: images, voice, and an avatar explaining the answer.","text":"AI video at Numerade Numerade Labs, Inc. At Numerade, we also built a way to make educational videos with AI: images, voice, and an avatar explaining the answer. AI Product & Engineering Lead Apr 2023–May 2024 ~1/5 Content cost 50+ Academic partners 90%+ Rollout adoption What I worked on We combined generated images, ElevenLabs speech, and HeyGen avatars to create videos near-instantly, at about a fifth of the cost of educator-created videos. We rolled it out across more than 50 academic partners. Adoption within that rollout passed 90%. Tools used ElevenLabs HeyGen More work"},{"id":"spectrum","title":"Making live conversations safer","kind":"work","url":"https://www.piyushvyas.com/work/spectrum/","summary":"I started with a dataset and a question: could we detect harmful speech directly from audio? That became a voice moderation product, first in English, then in 12+ languages.","text":"Making live conversations safer Spectrum Labs, Inc. I started with a dataset and a question: could we detect harmful speech directly from audio? That became a voice moderation product, first in English, then in 12+ languages. Senior Machine Learning Scientist Oct 2021–Apr 2023 12+ Languages Voice + text Technical leadership ONNX Model deployment What I worked on I built the labeled audio dataset and trained the first voice-native models. I worked with the CTO and engineers to turn them into a product customers could use. Later, I took on text moderation too: moving from classical classifiers to transformers, training on our data, and deploying with ONNX. I also joined customer and sales calls to understand what people needed. Tools used ONNX More work"},{"id":"comcast","title":"Speech recognition, with less error","kind":"work","url":"https://www.piyushvyas.com/work/comcast/","summary":"I worked on speech recognition for Comcast’s voice products. I also coauthored a paper about letting a model stop listening once it knows the command.","text":"Speech recognition, with less error Comcast Applied AI I worked on speech recognition for Comcast’s voice products. I also coauthored a paper about letting a model stop listening once it knows the command. Research Intern (NLP) Aug 2021–Oct 2021 30% Lower word error rate 2022 IEEE ICASSP wav2vec 2.0 Speech architecture What I worked on I built a wav2vec 2.0 speech-recognition system that reduced word error rate by 30%. Our separate ICASSP 2022 paper explored early exits in streaming command recognition. The full paper is below. Streaming speech recognition IEEE ICASSP 2022 · PDF · 5 pages More work"},{"id":"research","title":"Papers & awards","kind":"information","url":"https://www.piyushvyas.com/research/","summary":"Piyush Vyas’s coauthored research on speech translation, speech assessment and streaming recognition, including papers and award attribution.","text":"Papers & awards Speech translation, quality & recognition Speech translation INTERSPEECH 2021 · PDF · 5 pages Speech assessment IEEE ICASSP 2021 · PDF · 5 pages Streaming speech recognition IEEE ICASSP 2022 · PDF · 5 pages"},{"id":"speech-translation","title":"Speech translation","kind":"research","url":"https://www.piyushvyas.com/research/speech-translation/","summary":"We looked at how to give a Transformer useful structure for translating speech directly into another language.","text":"Speech translation INTERSPEECH 2021 Optimally Encoding Inductive Biases into the Transformer Improves End-to-End Speech Translation Piyush Vyas, Anastasia Kuznetsova, Donald S. Williamson 🏆 Best Student Paper Award We looked at how to give a Transformer useful structure for translating speech directly into another language. Open PDF to zoom or download 1 of 5 2 of 5 3 of 5 4 of 5 5 of 5 All papers"},{"id":"speech-assessment","title":"Speech assessment","kind":"research","url":"https://www.piyushvyas.com/research/speech-assessment/","summary":"Can a model judge speech quality without a clean reference recording? We trained one to estimate several measures of quality at once.","text":"Speech assessment IEEE ICASSP 2021 An End-to-End Non-Intrusive Model for Subjective and Objective Real-World Speech Assessment Using a Multi-Task Framework Zhuohuang Zhang, Piyush Vyas, Xuan Dong, Donald S. Williamson 🏆 Outstanding Student Paper Award Can a model judge speech quality without a clean reference recording? We trained one to estimate several measures of quality at once. Open PDF to zoom or download 1 of 5 2 of 5 3 of 5 4 of 5 5 of 5 All papers"},{"id":"streaming-speech","title":"Streaming speech recognition","kind":"research","url":"https://www.piyushvyas.com/research/streaming-speech/","summary":"We explored recognizing a spoken command before the speaker finishes, so the system can respond sooner.","text":"Streaming speech recognition IEEE ICASSP 2022 Temporal Early Exiting for Streaming Speech Commands Recognition Raphael Tang, Karun Kumar, Ji Xin, Piyush Vyas, Wenyan Li, Gefei Yang, Yajie Mao, Craig Murray, Jimmy Lin We explored recognizing a spoken command before the speaker finishes, so the system can respond sooner. Open PDF to zoom or download 1 of 5 2 of 5 3 of 5 4 of 5 5 of 5 All papers"},{"id":"tiny-parking-agent","title":"How I Would Design a Tiny Parking Agent for Visitors","kind":"writing","url":"https://www.piyushvyas.com/blog/tiny-parking-agent/","summary":"A design concept for an iMessage-to-browser parking loop, including failure states, privacy, and approval boundaries.","text":"How I Would Design a Tiny Parking Agent for Visitors Design concept · 4 min The parking portal for my building is simple, but it is still a task: open the site, enter the visitor's plate, choose the state, pick the duration, submit, and send back confirmation. That is exactly the kind of small, annoying workflow I like giving to an agent. The happy path should feel like texting a picture, not learning a parking system. The version I would build starts with a visitor sending me a rear-car photo over iMessage. The plate has to be visible. The agent reads the image, extracts the plate, asks for missing details, opens the parking portal in a browser worker, fills the form, submits it, and replies with the result. The key design choice is that the model does not get to improvise the whole job. It gets narrow tasks: read the plate, normalize the state, decide whether confidence is high enough, and write a short message. The browser worker owns the actual portal interaction. photo arrives extract plate and confidence ask for missing fields if needed open portal in isolated browser submit only after policy checks pass send confirmation or next question Most of the engineering is in the edge cases. A useful home agent needs a boring fallback for every confident-looking failure. If the plate is blurry, it asks for another photo. If the state is missing, it asks one question. If two plates are visible, it stops. If the portal says the vehicle is already registered, it sends that as the confirmation instead of submitting again. If the portal changes layout, the browser worker should fail closed and ask me to review. I would also separate unknown visitors from trusted ones. An allowlisted sender can register their own car within normal limits. A new sender triggers a one-time owner approval: \"Register ABC123 for Maya for tonight?\" That keeps the social workflow easy without letting any random message become a portal action. The always-on part is less glamorous. iMessage makes the intake awkward because the cleanest personal setup is still an always-on Mac signed into a dedicated account. A small watcher can turn incoming messages into jobs, and a separate worker can handle OCR, policy, and browser automation. The private parts are separated from the browser worker, and photos do not need to live forever. For 24/7 reliability, I would keep the system small: a Mac mini or home server for message intake, a queue for retries, a browser worker that can restart cleanly, and a health check that alerts me when the portal login or watcher breaks. Portal credentials belong in a secrets vault, not in prompts or logs. The guardrails matter more than the automation. Rate limit registrations. Redact plates in long-term logs. Delete photos after a short window. Keep screenshots only when a run fails. Require manual approval for low-confidence OCR, unknown senders, paid actions, or anything that changes account settings. This is not an enterprise platform. It is a home project. But the lesson is the same as larger agent systems: the model is only one piece. The product is the loop around it, the permissions it respects, and the way it behaves when the world is messy. All writing"},{"id":"harness-engineering","title":"Harness Engineering","kind":"writing","url":"https://www.piyushvyas.com/blog/harness-engineering/","summary":"The reliability work around an agent: context, tools, permissions, sandboxing, memory, and traces.","text":"Harness Engineering Agent reliability · 3 min When an agent fails, the model gets blamed first. Sometimes that is fair. More often, the failure is in the harness: the code around the model that decides what it sees, what it can call, what gets checked, and when the loop stops. The model is only one box. The harness decides how that box touches the real world. I use \"harness\" for the runtime frame around an agent. It includes the system prompt, context engine, tool interface, memory, permission checks, sandbox, loop controller, subagent routing, and observability. The distinction I care about is this: scaffolding is what you set up before the run. Harness is what executes every turn. assemble context call model parse tool call check policy run tool record result decide whether to continue Most production failures I have seen live in three places. Context. The model only knows what the harness includes. If the original goal falls out of the window, the agent drifts. If every tool schema is loaded every turn, the context budget disappears before the work starts. Good harnesses treat context like a budget, not a transcript. Dispatch. The moment a tool runs is the moment the agent touches the world. Permissions, allowlists, approval gates, and audit logs belong here. A prompt that says \"do not edit backup files\" is weaker than a file tool that refuses .bak paths. Verification. Agents are good at saying work is done. The harness should check claims that matter. If the model says \"tests pass,\" the harness can run the test command or mark the claim as unverified. If the model says \"deployment succeeded,\" the harness can require the deploy artifact or log link. One concrete example: an agent was asked to fix auth.ts and edited auth.ts.bak instead. The model looked confident, and the diff looked plausible. The fix was not a better instruction. It was a dispatch policy: only editable source extensions, explicit path matching, and a unit test around the file tool. This is why I think the highest-leverage agent work is below the model. Narrow tools. Typed outputs. Real permissions. Sandboxed execution. Memory with provenance. Traces you can inspect. These are not glamorous pieces, but they decide whether an agent is a demo or something you can hand to a team. If you build agents, ask what the harness enforces. If you buy agents, ask to see the tool policy and logs. If you use agents, remember that a bad result is often not \"the AI being weird.\" It is usually a missing check around the AI. All writing"},{"id":"introducing-skill-tools","title":"Deterministic Tooling for Agent Skills","kind":"writing","url":"https://www.piyushvyas.com/blog/introducing-skill-tools/","summary":"A private routing failure led me to build public, inspectable checks for SKILL.md structure, quality, and selection.","text":"Deterministic Tooling for Agent Skills Public tooling · 6 min I asked a skill router to “build an MCP server.” It returned a Figma skill. BM25 had not malfunctioned. The Figma skill mentioned “MCP server,” and the intended skill described its job with different words. The router matched the text I gave it. That private failure changed how I think about SKILL.md : its description is not introductory copy. It is runtime metadata. The failure came from a private audit of 53 skills in December 2025. That corpus and its raw results are not public, so the count and the Figma misroute are an origin story, not a benchmark. What is public is the response: skill-tools , a TypeScript toolchain for inspecting agent skills without asking another model to grade them. A skill can fail before the model reasons One of the least glamorous defects can become expensive in practice. A skill tells the agent to read references/policy.md , but the file is missing. The agent can fail on the path, consume context recovering, or improvise around the missing instruction. The public parser catches that class of defect before runtime. It checks the frontmatter, required name and description, Markdown body, directory naming, and references under scripts/ , references/ , and assets/ . It also reports advisory token and line budgets. The linter adds visible rules for problems such as hardcoded machine paths, secret-like text, and vague trigger language. npx skill-tools check ./skills error file-reference-exists references/policy.md warn no-hardcoded-paths /Users/me/project warn specific-description \"Manage deployments\" These checks are intentionally boring. At the same version and configuration, the same file should produce the same diagnostics in a terminal, pre-commit hook, GitHub Action, SARIF report, or browser session. The score is a rubric, not a truth machine skill-tools exposes a 0–100 score across five named dimensions: description quality, instruction clarity, specification compliance, progressive disclosure, and security. Each finding is inspectable. That makes the score a useful review signal and CI guardrail. It does not prove that a skill will improve an agent’s task success. I have not published a benchmark showing that a 90-point skill outperforms a 70-point skill in downstream work. The honest interpretation is narrower: the score summarizes a visible static rubric and makes regressions easy to detect. Deterministic routing is explainable—not infallible The default public router uses BM25. It indexes a skill’s description plus bounded context extracted from its body and section headings. With a configured threshold, it can return no result, which gives the caller a practical abstention path; the default threshold is zero. The library also exposes optional embedding-provider interfaces, but BM25 remains the inspectable, zero-external-dependency default. I like that default because a bad result can be debugged in ordinary language: which terms matched, which useful terms were absent, and whether two skill descriptions overlap. The fix may be a better description, better indexed context, a higher threshold, or a reranker. “Deterministic” only means the same inputs, version, and configuration produce the same choice. It does not mean the choice is correct. That distinction is the lesson from the Figma route. A hidden model judgment would have made the error harder to explain. A term-based result made the weak contract visible. One implementation, several enforcement points The public project packages the same core behavior for different moments in a skill’s life: authoring: watch mode and a pre-commit hook shorten the feedback loop; review: CLI output and a five-part score make the change legible; CI: a GitHub Action and SARIF can block known structural defects; generation: OpenAPI and MCP schemas can produce a draft skill that passes through the same checks; and runtime: the router can select, rank, filter, or return no match when nothing clears a chosen threshold. The browser playground states that its tools run locally in the page without an LLM call, API call, or file upload; that behavior can also be inspected in the browser. The choice is part of the product, not a novelty. Skills often contain internal paths, operating procedures, and occasionally a secret pasted where it should not be. A privacy checker should not require uploading the file to someone else’s server. What remains to prove The repository publishes the parser, checks, lint rules, scoring logic, router, generators, tests, packages, CI workflow, and release history. The playground lets a visitor exercise much of that behavior directly. Those are independently inspectable product claims. The next useful evidence is not another feature. It is a revision-pinned public benchmark: a published skill corpus, a realistic query set, expected routes, top-k results, false positives, abstentions, and rerun commands—followed by task-level evals after selection. Then I can test whether a lint or routing change improves actual agent outcomes instead of merely improving its own score. Evidence note: the implementation, public CI, releases, and local browser playground are independently checkable. The original 53-skill audit and Figma misroute are self-reported. No published result yet establishes score-to-task-success correlation or routing superiority. Source / Documentation / Try locally in the browser All writing"},{"id":"browser-agents","title":"Booking Open Play Before It’s Gone","kind":"writing","url":"https://www.piyushvyas.com/blog/browser-agents/","summary":"The Thursday-noon reservation race that taught me where scheduling ends, browser semantics begin, and human approval belongs.","text":"Booking Open Play Before It’s Gone Browser agents · 5 min Every Thursday at noon, open-play slots would drop. The useful Tuesday session was often gone before I finished the same ridiculous routine: refresh, filter, tap, hope. That is not a workflow. It is a thumb race. From my private notes, I got the slot twice in six weeks manually. So I built a small browser agent around one narrow job: watch the release window, find a session that matched my rules, prepare the registration, and ask me before the paid confirmation. The job was smaller than “book pickleball” A vague booking objective hides several different responsibilities: Schedule: trigger at the release window instead of asking an agent to wait inside a session. Observe: read the page’s controls, labels, availability, and current price through the DOM. Decide: match only the day, time, level, club, and price range I had already specified. Approve: stop before the action that created a charge. Verify: require a reservation identifier or explicit success state; otherwise report failure. Making those boundaries explicit turned a flashy browser demo into a routine I could reason about. The conversation needed one useful question Agent: Tuesday, 6:30 PM, intermediate, 4 spots, $12. This matches your rules. Book it? Me: Yes. Agent: Reserved. Confirmation saved. I did not want the model negotiating preferences while inventory disappeared. The preferences were already data. The only decision that still needed me was the consequential one. DOM first, pixels only if the page forces it This was an internet task with useful page semantics, so structured browser automation was the right interface. The agent targeted controls by role and accessible name, read availability as page state, filled the form, and checked the confirmation surface. A vision model hunting for the “Register” button in screenshots would have added ambiguity without adding value. The selector also had to survive small page changes. I preferred meaning—“Register,” the session time, the level label—over an opaque CSS path. If a required control disappeared or the session details changed, the run stopped instead of guessing. The real race condition was between observe and submit A slot can disappear after the page is read but before the registration is submitted. Price or session details can also change. That means approval cannot freeze stale state forever. Immediately before submission, the worker reacquired the intended session and rechecked the fields that mattered. A mismatch invalidated the approval and returned a clear failure. The safe behavior was not “try harder.” It was “this is no longer the action you approved.” What the field test told me My private run log recorded browser work at about four seconds on my setup, versus a manual path usually around fifteen seconds. During a four-week test, the recorded runs selected the intended session. The exact run count, raw log, and reservation artifacts are not public, so those numbers are directional—not a benchmark. The more durable result was the design: a deterministic schedule instead of an agent waiting around; structured page observation instead of pixel hunting; predeclared preferences instead of model improvisation; fresh-state validation immediately before the consequential action; human confirmation before payment; a receipt or explicit failure instead of a narrated success. That pattern extends beyond pickleball. It fits appointment windows, limited inventory, recurring form submissions, and other small internet tasks where speed matters but an incorrect action matters more. What came later This October 2025 field test predates the public Browser Agent Protocol repository , which was created in February 2026. BAP is related implementation evidence for the semantic browser layer I subsequently published; it is not a public record of these reservation runs. Evidence note: The workflow, timing, and four-week outcome are reconstructed from my private notes and are not independently reproducible. No credentials, reservation receipts, or private portal details are published. Treat this as a personal field report and a design lesson—not a performance claim for BAP. All writing"},{"id":"introducing-browser-agent-protocol","title":"Browser Agent Protocol","kind":"writing","url":"https://www.piyushvyas.com/blog/introducing-browser-agent-protocol/","summary":"A protocol layer for browser agents: semantic observation, structured actions, and one browser interface.","text":"Browser Agent Protocol From the archive · November 9, 2025 The first browser-agent bug I kept seeing was not an AI bug. It was a selector bug. The agent clicked .auth-actions button:first-child . After a redesign, that selector pointed at \"Sign Up\" instead of \"Sign In.\" The agent did exactly what the code told it to do. BAP targets page meaning first: role, name, label, and state. CSS is the escape hatch. Browser Agent Protocol is my attempt to make that boundary explicit. An agent should not need a custom Playwright wrapper for every project. It should get a small set of browser tools: navigate, observe, click, fill, extract. The observation should be structured enough for a model to reason over, and stable enough to survive normal UI churn. The core move is semantic selectors: // Fragile: where the button happened to be await page.click('.auth-form button[type=\"submit\"]') // BAP-style: what the button is await browser.click({ role: 'button', name: 'Sign In' }) The accessibility tree already gives browsers this information. A screen reader does not care that a button uses Tailwind classes or sits inside three wrapper divs. It cares that the element is a button named \"Sign In.\" Browser agents need the same abstraction. BAP exposes that through an MCP server and CLI. A typical turn is: observe the page, choose a semantic target, act, then observe again. The refs are stable inside the session, and the action log gives you something to debug when a page behaves differently than expected. What I like about this design is that it does not pretend the web is clean. Pages have modals, iframes, hidden buttons, slow network calls, and inconsistent markup. BAP does not make those disappear. It gives the agent a narrower, more inspectable interface to work through. The repo is public here: browseragentprotocol/bap . All writing"},{"id":"introducing-useid","title":"Universal Semantic Element ID","kind":"writing","url":"https://www.piyushvyas.com/blog/introducing-useid/","summary":"A portable element signature for finding the same browser control across runs.","text":"Universal Semantic Element ID From the archive · March 12, 2026 BAP can give an agent a stable ref inside one browser session, like @e4 . Close the browser and that ref is gone. uSEID is the layer I built for the next question: how does an agent find the same element tomorrow? uSEID combines meaning, DOM context, and position, then resolves with a confidence score. The problem shows up quickly. A page has three \"Submit\" buttons. Or the \"Sign In\" label changes to \"Log In.\" Or a deploy inserts a \"Sign Up\" button above the old target. CSS and session refs do not give the agent enough identity to say, \"this same element as before.\" uSEID builds a signature from three signals: Signal What it captures Semantic ARIA role and accessible name, like button \"Sign In\". Structural Ancestor roles, nearby labels, sibling context, and depth. Spatial Where the element appears in the viewport. Build phase: pass DOM and accessibility snapshots plus the target element. uSEID returns a JSON signature you can store. Resolve phase: pass the stored signature and the current page snapshot. uSEID scores candidates, returns the best match, or abstains when the match is too weak or ambiguous. build signature today deploy changes the page resolve signature tomorrow return match or abstain The abstention behavior is the part I care about most. In browser-agent workflows, clicking the wrong element is usually worse than stopping. uSEID checks origin and path binding, role compatibility, confidence, and ambiguity before returning a target. This is useful for monitoring, regression checks, multi-session workflows, and agent memory. If an agent learns that a specific control starts checkout, I want that memory to be a portable reference, not a brittle CSS path. GitHub / npm All writing"},{"id":"deterministic-browser-agents","title":"I Replayed a Browser-Use Agent Session Deterministically","kind":"writing","url":"https://www.piyushvyas.com/blog/deterministic-browser-agents/","summary":"A browser-use run captured with DBAR and replayed later with matching strict observables.","text":"I Replayed a Browser-Use Agent Session Deterministically From the archive · March 26, 2026 Browser agents are easy to demo and hard to audit. The question I wanted to answer was narrow: can I record one real browser-use session at the browser boundary and replay it later with matching strict observables? The experiment: capture a session, store a capsule, replay it, and compare strict hashes. The setup used two processes. The Python process ran browser-use against a Chromium instance. A Node process attached to the same browser over CDP and recorded the session with DBAR. browser-use agent | v Chromium over CDP | v DBAR capture -> capsule -> replay DBAR recorded virtual time, network requests and responses, DOM snapshots, accessibility snapshots, screenshots, and hashes at step boundaries. Twenty-four hours later, I replayed the capsule on a fresh browser instance. For this session, the strict observables matched: DOM hash, accessibility-tree hash, and network digest. That is the important phrasing. I am not claiming every browser session can be replayed perfectly. Fonts, GPU differences, cross-origin behavior, service workers, browser versions, and unsupported APIs can all introduce drift. I am saying this captured Chromium session replayed cleanly on the observables DBAR treats as strict. The more useful case is divergence. If a replay fails, DBAR reports the step and observable: step: 3 observable: dom expected: a3f2... actual: b7c1... That changes the debugging loop. Instead of rerunning and hoping the failure appears again, you inspect the recorded step where the page stopped matching the capsule. For agent systems, this is the kind of evidence I want: not just \"the agent clicked submit,\" but the browser state before and after that click, plus the network data that got it there. Logs tell you what the harness thought happened. Replay evidence tells you what the browser exposed at the time. DBAR is public here: github.com/pyyush/dbar . All writing"}]}