What 2,500 Hackathon Entries Taught Me About Spines, Sky Maps, and Agentic Leases

The verdict for the WebMCP Challenge is out: 7,000 registered souls, roughly 2,500 submissions from over a hundred countries, all hammered together across ten caffeine-soaked days.

I don’t envy the judging panel. Staring at 2,500 competing agentic architectures is the sort of cognitive punishment usually reserved for low-level tech purgatory.

My entry, DealTable, did not place in the top ten. And genuinely? That is completely fine. It was feral, slightly unhinged to build, and I’m thoroughly glad it exists.

The Pitch for the Fast-Scrolling Class

DealTable is a negotiation clean-room for the coming agentic web.

You hand it a mandate: your target, your walkaway threshold, and precisely how far you’re prepared to bend before dignity gives up. A counterparty pushes back in the chat. ChatGPT sits in the corner, acting as your advisor by calling typed WebMCP tools directly on the live page—not by hallucinating UI elements or squinting at raw DOM trees like an exhausted temp at 3:00 AM.

The moment a deal crosses your walkaway line, Human-in-the-Loop trips the emergency brake. Everything worth dragging a counterparty into court over gets stamped into an audit trail.

I built it because the real dystopia lurking in our automated future isn’t a gleaming chrome skull crushing a ribcage; it’s an overly polite corporate agent cheerfully agreeing to contractual terms that quietly turn your mortgage into a multi-generational performance art piece.

I just wanted a room where the human still has a hand firmly on the shock collar.

A man with a beard and glasses sits at a wooden table holding a cup, next to an old computer displaying green text on the screen. The setting is dimly lit with technological equipment in the background.

The Gallery as a Free Masterclass

After the shutters came down, I went prowling through the project gallery. Some of the builds were legitimately intimidating: spatial editors, shared notebooks, medical viewers, and collaborative wedding seating charts where human and model stare down the same live artifact.

That is the high-water mark now.

Mandate taking the crown made absolute sense. They were playing in the exact same theoretical pool as DealTable—authority, boundaries, delegation, and who gets to touch what—but they took the mechanics a leap further. Instead of just slapping an apologetic “Click Approve to continue” prompt onto a static API list, the tool surface itself mutates dynamically as permissions are granted or stripped away. Authority expires, the tools evaporate.

That wasn’t just a win; it was an architectural lecture I sat in the front row for, taking furious notes.

My internal post-mortem was ruthlessly mundane:

  • The Pitch Video: Too many Scots-accented ums, not enough surgical ruthlessness. Lead in lockstep: Problem → Mandate → Tool Call → Human-in-the-Loop → Closure.
  • The Narrative: Lead with the visceral bleeding problem, not the plumbing. WebMCP is the pipeline; nobody falls in love with the pipe.
  • The Execution: V1 ran a rule-based counterparty so I could focus entirely on the advisor bridge. Honest engineering, but in a hackathon arena, the theatre needs to be agent-native from lobby to exit door.
A person leans against a stone pillar in a dimly lit, futuristic environment, surrounded by glowing holographic displays of celestial navigation, a scroll, a dungeon layout, and an adaptive floor plan. A sign on the wall includes a humorous quote about coffee.

Five Winners I’d Actually Boot Up (And One I Suspect Solved a Judge’s Mortgage Dilemma)

Two thousand five hundred entries generate an extraordinary amount of speculative noise. These are the five that sliced clean through it.

1. MASIL — The One That Actually Matters

Reconnecting Korean elders to creative life.

An agent suggests classic calligraphy references and surfaces Janggi moves via WebMCP, but leaves the brushwork and the completion of the game entirely to the elder.

This is the project I’d back in a noisy pub debate. Not because it had the most complex stack, but because the human stakes are palpable. Elders isolated at home, their creative universe contracting to the size of a glowing piece of glass. The agent isn’t stepping in to replace human volition; it acts as an errand boy to fetch the pieces of a world they can no longer easily walk out to, handing them back as tools for their own craft.

That is WebMCP with a moral compass: typed tools, a shared page, and basic human dignity preserved intact. It makes you feel slightly guilty that your own entry was built around defending freelance day rates.

2. Roque Nights — Built by Someone Who Looks at the Sky

A stargazing planner from an engineer at the Gran Telescopio Canarias.

The agent crunches seeing conditions, isolates the darkest nights, figures out what’s actually worth observing, and slews the interactive sky map into alignment while you sit back and watch.

Unreservedly brilliant. Zero slide-deck jargon about “agentic paradigm synergy.” Just an engineer who genuinely cares about photons, an automated assistant doing the tedious atmospheric arithmetic, and a shared visual map both are looking at. If you’ve ever stood in a freezing field at midnight wondering if it was worth hauling 20 kilograms of glass out of the boot of your car, you immediately understand why this exists.

3. Observatory — Unfortunate Name, Impeccable Dungeon Energy

A procedural map-maker for tabletop designers and weary dungeon masters.

Type “a lonely crumbled watchtower in a choking swamp,” and the system doesn’t spit out a hallucinatory jpeg with six fingers. It generates structured, fully editable vector terrain, rooms, and labels through WebMCP.

If teenage me had been handed this, I’d have treated the monitor like a sacred religious relic. The beauty isn’t “the machine drew a landscape.” The triumph is that it created a manipulable environment that you and your gaming group can spend four hours arguing over. Rename it so it doesn’t sound like a venture-backed SaaS platform from 2013, and ship it.

4. Alza — The Antidote to Cynicism

A floor-plan editor that surrenders its tools directly to the model.

Feed it a single snapshot and one verified dimension; it traces walls, extracts spatial dimensions, and constructs a navigable 3D model you can stroll through in your browser.

This is the demo you pull out when a non-technical sceptic asks why everyone is screaming about agentic standards. It’s structure over pixels. The agent handles the tedious measurement math; the human remains the architect of the living space.

It has a grounded utility that makes DealTable feel like a late-night corporate philosophy seminar—which, let’s be honest, it essentially was.

5. ArchMorph — And the “Someone on the Jury is Renovating” Hypothesis

Fifty-seven distinct architectural tools, dynamic floor plans, and layout validation.

A staggering build. It’s also remarkably close in theme to Alza. When two distinct entries in a top-ten roster both revolve around floor plans, structural walkthroughs, and spatial clearances, one cannot help but suspect at least one member of the judging panel was mid-kitchen-extension and having an absolute nightmare with their contractor.

I’m not levelling accusations. I’m just saying if an investigative committee looked into it, I wouldn’t be surprised.

What Winning Actually Looked Like

The judges didn’t care about developers merely chanting “look, we registered seven API endpoints.”

They rewarded a single, shared artifact on screen—a celestial chart, a battle map, an architectural canvas, a digital calligraphy desk—and a division of labour where the machine acts as the grease and the human remains the spine.

DealTable v2 is already queued up on the local machine. The spine stays; the execution gets sharpened with the lessons pulled straight from the winners’ circle.

For the curious or the brave:

A massive tip of the hat to the WebMCP organisers (OpenAI and DevPost), the public builders, and my long-suffering collaborative stack that helped me figure out what needsHitl means in anger (including that brief, glorious window where DealTable accidentally attempted to negotiate commercial real estate at $undefined/sqft).

Huge congratulations to the winners—particularly Mandate. You set the bar.

The Day I Outsourced My Backbone to an MCP Bridge

There is a distinct, uniquely contemporary brand of nausea reserved for corporate negotiation.

It usually begins on a Tuesday afternoon. You are sitting in the grey light of a laptop screen, staring at an email from an enterprise client whose signature line is longer than the Magna Carta. They want a thirty percent discount on your day rate. They are also proposing a payment schedule calibrated to conclude somewhere near the heat death of the universe.

Traditionally, you do what any self-respecting contractor does: you stare into the middle distance, do feverish mental arithmetic to calculate if you can survive on dry pasta until November, and draft a reply dripping with performative corporate politeness. You write, “Thanks for reaching out! Happy to find a middle ground,” while your spleen violently retracts into your ribcage.

The horror isn’t the money. The horror is the slow, wet erosion of your dignity in an unmonitored thread with no audit trail.

So, when the WebMCP hackathon opened with Devpost, I didn’t see an emerging browser protocol. I saw a containment unit. I decided to build DealTable: a clean, sterile room where synthetic entities could barter for scraps of my mortal labour, strictly supervised by a digital shock collar.

The Mathematics of Keeping Your Spine

In modern software architecture, people are terrified of artificial intelligence turning rogue, seizing missile silos, and sterilising the biosphere.

Personally, I am far more terrified of an AI agent negotiating a vendor contract on my behalf and cheerfully agreeing to a 90-day Net payment term because its sentiment-analysis model mistook predatory procurement tactics for “a collaborative synergy opportunity.”

To prevent the machine from liquidating my mortgage in the name of algorithmic politeness, DealTable relies on an ancient, barbaric concept: The Mandate.

Before you let the silicon speak, you set your floor. The target price. The concession tolerance. And, crucially, the walkaway limit ($w$).

The governance model is brutal:

$$\operatorname{apply}(p) = \begin{cases} \text{execute immediately},
& \text{if } p \ge w \\ \text{require human approval}, & \text{if } p < w \end{cases}$$

If the counterparty proposes an offer above your floor, the machines trade pleasantries and execute. The moment an offer dips even a fraction of a penny below your survivable threshold, execution freezes. The machine stops dead. It turns its digital head, looks you dead in the eyes, and demands an explicit, verified click: Approve or Reject.

Agents may hallucinate poetry, optimize supply chains, or pretend they have souls. But they cannot cross the floor. The human still owns the boundary.

Inside the Terrarium: How DealTable Operates

I built DealTable on a lean, serverless spine: Vite, React, and Tailwind, hosted on Vercel with local session state. No heavyweight databases or enterprise middleware—just pure client-side orchestration so hackathon judges could witness algorithmic bartering without signing their lives away to an authentication provider.

The secret sauce isn’t prompt engineering. It’s the Model Context Protocol (MCP).

Instead of treating an LLM like an omniscient wizard squinting at screenshots through brittle DOM automation, DealTable turns the browser window into a strictly typed operating theater. The page registers seven dedicated WebMCP tools:

  • get_deal_state: Ingests the current mandate, active asking price, conversation history, and live metrics.
  • set_mandate: Reconfigures the operational boundaries when market conditions sour.
  • parse_opening_offer: Strips incoming corporate jargon down to cold, quantifiable digits.
  • propose_offer & concede: Executes calculated counter-punches within authorized bounds.
  • hold_firm: An programmatic, polite equivalent of a flat refusal.
  • accept_term: Executes the closing sequence behind a cryptographic gate.

When ChatGPT acts as the advisor, it isn’t guessing. It is pulling live, validated state from the page and calling specific, sandboxed functions. When an action threatens the walkaway limit, the tool returns needsHitl.

A high-contrast Human-in-the-Loop modal slams over the viewport. The underlying JavaScript promise refuses to resolve until an actual warm-blooded creature clicks a button. The advisor is held in suspended animation, trapped in execution limbo while the human decides if the insult is tolerable.

Blood on the Terminal

Building autonomous negotiation arenas sounds pristine until you actually wire the pipes and watch the plumbing back up.

First came the bureaucratic indignities. Vercel threw an existential fit because a project repository contained capital letters and whitespace. A rogue hash mark sitting inside an npm run build command quietly sabotaged the Vite bundle like a loose bolt dropped into an aircraft turbine on deadline day.

Then came the existential bugs. During early test runs, our HITL modal inadvertently ingested an unparsed response object instead of the clean pending price structure. The result? A triumphant, unhinged bot proudly locking in a commercial lease for precisely $undefined/sqft. A dystopian victory for zero-cost real estate, perhaps, but tricky to defend in an audit.

Worse, demoing the workflow in real time created a bizarre psychological standoff. Triggering HITL via an active WebMCP tool call paused ChatGPT indefinitely, waiting for a human click on the host page. On a three-minute video recording, an AI staring wordlessly into space looks less like cutting-edge governance and more like catastrophic system failure. We restructured the demo rail to demonstrate deterministic tool execution in parallel with manual override triggers—making the safety rails obvious without subjecting the judges to awkward digital silence.

What the Silence Taught Me

We emerged from the hackathon with a working URL, a live audit export, and a few stark truths about the coming synthetic economy:

  1. Typed tools beat automated vision every single time. Letting an LLM scrape a webpage to make financial decisions is digital negligence. Forcing it to call strictly typed, validated tools that mutate isolated state is the only way to retain sanity.
  2. Never give one bot two jobs. In our early drafts, a single LLM tried to play both the cutthroat vendor and the impartial advisor. It rapidly degenerated into a schizophrenic pantomime where the model essentially negotiated with its own hallucinations. You must separate the actors: the principal sets the mandate, the counterparty pushes their agenda, and the external advisor sits outside the transaction.
  3. The safety switch cannot hide in a sub-menu. If human oversight is buried three clicks deep or masked behind opaque JSON logs, it doesn’t exist. HITL belongs front and centre—a flashing, unavoidable perimeter wire.

The Road to the Silicon Souk

The prototype works, but the future is significantly weirder.

Next comes swapping out our deterministic, rule-based adversary for a fully conditioned LLM adversary—one capable of simulating specific, predatory procurement personas (the Passive-Aggressive Startup Founder, the Enterprise Bureaucrat with Infinite Runway, the Venture-Backed Lowballer). After that, automatic parsing of 80-page commercial PDF contracts, stripping away the legalese to find the hidden clauses that usually bite you six months later.

Ultimately, we are barrelling toward an internet where autonomous agents will spend their days aggressively bartering with other autonomous agents over micro-transactions, service level agreements, and server runtime fees.

If we don’t build deterministic floors into the code now, we will wake up in a decade to discover our synthetic representatives have cheerfully traded away our rights, our margins, and our weekends—all to achieve a 98% polite closure metric.

I’d rather keep the walkaway limit in React state, thanks. At least when the world ends, my console will log the exact price at which I refused to sell out.

#Devpost #DealTable #WebMCP #Vercel #netlify

Apple and Google: A Forbidden Love Story, with AI as the Matchmaker

Well, butter my biscuits and call me surprised! Apple, the company that practically invented the walled garden, has just invited Google, its long-standing frenemy, over for a playdate. And not just any playdate – an AI-powered, privacy-focused, game-changing kind of playdate.

Remember when Apple cozied up to OpenAI, and everyone assumed ChatGPT was going to be the belle of the Siri-ball? Turns out, Apple was playing the field, secretly testing both ChatGPT and Google’s Gemini AI. And guess who stole the show? Yep, Gemini. Apparently, it’s better at whispering sweet nothings into Siri’s ear, taking notes like a diligent personal assistant, and generally being the brains of the operation.

So, what’s in it for these tech titans?

Apple’s Angle:

  • Supercharged Siri: Let’s face it, Siri’s been needing a brain transplant for a while now. Gemini could be the upgrade that finally makes her a worthy contender against Alexa and Google Assistant.
  • Privacy Prowess: By keeping Gemini on-device, Apple reinforces its commitment to privacy, a major selling point for its users.
  • Strategic Power Play: This move gives Apple leverage in the AI game, potentially attracting developers eager to build for a platform with cutting-edge AI capabilities.

Google’s Gains:

  • iPhone Invasion: Millions of iPhones suddenly become potential Gemini playgrounds. That’s a massive user base for Google to tap into.
  • AI Dominance: This partnership solidifies Google’s position as a leader in the AI space, showing that even its rivals recognize the power of Gemini.
  • Data Goldmine (Maybe?): While Apple insists on on-device processing, Google might still glean valuable insights from anonymized usage patterns.

The Bigger Picture:

This unexpected alliance could shake up the entire tech landscape. Imagine a world where your iPhone understands your needs before you even ask, where your notes practically write themselves, and where privacy isn’t an afterthought but a core feature.

But let’s not get ahead of ourselves. There are still questions to be answered. How will this impact Apple’s relationship with OpenAI? Will Google play nice with Apple’s walled garden? And most importantly, will Siri finally stop misinterpreting our requests for pizza as a desire to hear the mating call of a Peruvian tree frog?

Only time will tell. But one thing’s for sure: this Apple-Google AI mashup is a plot twist no one saw coming. And it’s going to be a wild ride.

Using OpenAI’s API

I enrolled in this course in May, a time when access to OpenAI was limited and its commercial model was still under development. Hence, leveraging the API emerged as the most straightforward method to use the platform. Jose Portilla’s course on Udemy brilliantly introduces how to tap into the API, harnessing the prowess of OpenAI to craft intelligent Python-driven applications.

The influx of AI platforms and services last summer indicates that embedding AI models into developments has become a standard practice.

OpenAI’s API ranks among the most sophisticated artificial intelligence platforms today, offering a spectrum of capabilities, from natural language processing to computer vision. Using this API, developers can craft applications capable of understanding and interacting with human language, generating coherent text, performing sentiment analysis, and much more.

The course initiates with a rundown of the OpenAI API basics, including account and access key setup using Python. Following this, learners embark on ten diverse projects, which include:

  • NLP to SQL: Here, you construct a POC that enables individuals to engage with a cached database and fetch details without any SQL knowledge.
  • Exam Creator: This involves the automated generation of a multiple-choice quiz, complete with an answer sheet and scoring mechanism. The focus here is on honing prompt engineering skills to format text outputs efficiently.
  • Automatic Recipe Creator: Based on user-input ingredients, this tool recommends recipes, complemented with DALLE-2 generated imagery of the finished dish. This module particularly emphasizes understanding the various models as participants engage with the Completion API and Image API.
  • Automatic Blog Post Creator: This enlightening module teaches integration of the OpenAI API with a live webpage via GitHub Pages.
  • Sentiment Analysis Exercise: By sourcing posts from Reddit and employing the Completion API, students assess the sentiment of the content. Notably, many news platforms seem to block such practices, labeling them as “scraping.”
  • Auto Code Explainer: Though I now use Co-pilot daily, this module introduced me to the Codex model. It’s adept at crafting docstrings for Python functions, ensuring that every .py file returns with comprehensive docstrings.
  • Translation Project: This module skims news from foreign languages, providing a concise English summary. A notable observation is the current model’s propensity to translate only to English. Users must also ensure they’re not infringing on site restrictions.
  • Chat-bot Fine-tuning: This pivotal tutorial unveils how one can refine existing models using specific datasets, enhancing output quality. By focusing on reducing token counts, learners gain insight into training data pricing, model utility, and cost-effectiveness. The module also underscores the rapid evolution of available models, urging students to consult OpenAI’s official documentation for the most recent updates.
  • Text Embedding: This segment was a challenge, mainly due to the intricate processes of converting text to N-dimensional vectors and understanding cosine similarity measurements. However, the module proficiently guides through concepts like search, clustering, and recommendations. It even delves into the amusing phenomenon of “model hallucination” and offers strategies to counteract it via prompt engineering.
  • General Overview & The Whisper API: Concluding the course, these tutorials provide a holistic understanding of the OpenAI API and its history, along with an introduction to the Whisper API, a tool adept at converting speech to text.

It’s noteworthy that most of the course material utilized the ChatGPT-3.5 model. However, recent updates have introduced a more efficient -turbo model. Additional information can be found here.

The course adopts a project-centric approach, with each segment potentially forming the cornerstone of a startup idea. Given the surge in AI startups, one wonders if this course inspired some of them.

This journey unraveled the intricate “magic” and “engineering” behind AI, emphasizing the importance of prompt formulation. Participants grasp essential elements like API authentication, making API calls, and processing results. By the course’s conclusion, you’re equipped to employ the OpenAI API to develop AI-integrated solutions. Prior Python knowledge can be advantageous.