Works

AI Chat Arena – The Digital Shrine 2025/03 ~ 2025/05

Gamifying RLHF through the Taiwanese Ritual of Zhijiao

Problem

Global LLMs lack Taiwanese cultural nuance, and the RLHF labeling needed to fix that is too tedious to sustain volunteers.

What I did

Product Designer. Used a ≈600-person interest survey to redirect the product from 'better chatbot' to 'Digital Oracle,' then designed the blind-comparison Arena and the Zhijiao ritual layer around a four-outcome label schema the model developer fixed in advance.

Evidence

Interest survey, n≈600, on which topics people would answer AI questions about for free — the finding that redirected the product from a better chatbot toward a Digital Oracle.

Status

Shipped as a working product, then the program was discontinued. The ritual metaphor remains a design hypothesis, never tested against a plain thumbs-up/down control.

Challenge

Global LLMs lack Taiwanese cultural nuance, and traditional methods for gathering human feedback (data labeling) are too tedious to sustain user engagement. While models like DeepSeek from China are powerful, they often lack the cultural nuance, political context, and Traditional Chinese vocabulary required for Taiwan's industries.

We needed to fine-tune these models for local usage. To fine-tune effectively, we needed human feedback (RLHF: Reinforcement Learning from Human Feedback), but asking volunteers to rate hundreds of text responses is boring and tedious.

Our objective: democratize local model training by building a gamified "Arena" that transforms complex data evaluation into an engaging, culturally resonant social activity.

Role

Product Designer

Timeline

2025.03 ~ 2025.05

Team

Master's-student teammates, a model developer, and a faculty advisor — sponsored by the Open Source Foundation in Taiwan

Skills

Interest Survey (n≈600) · Information Architecture · UI/UX · Figma · RLHF Workflow Design

How the team worked. Working alongside the model developer is what kept this from being a skin over a research problem. Because they could tell me what the training pipeline could actually consume, the interaction design and the label schema had to be settled together early, before any ritual framing was designed. My advisor and the other master's students pushed on the cultural framing, which is where the Zhijiao metaphor came from.

Live DemoView Interactive Prototype

Chat and Social Mode are live in the demo. Mission Mode was never built, and Social Mode's ranking layer runs on static sample data.

Approach

To identify the drivers that would motivate users to voluntarily train our AI, we surveyed approximately 600 people, asking which topics they would be willing to answer AI questions about for free. Generic technical prompts drew the least interest. What respondents were most willing to answer was hyper-local — food recommendations, neighborhood issues, and relationship advice, the last of these specifically appealing to 'Love Gods.'

This data revealed a critical strategic insight: users were not seeking a 'better ChatGPT,' but rather a 'Digital Oracle.' The standard, sterile chat interface was fundamentally ill-suited for the intimate, culturally specific conversations users actually cared about, necessitating a pivot toward a more ritualistic design.

The Design Strategy

Our goal was to create an "Arena" website interface where users blind-test two models (Model A vs. Model B) and vote on the better response. We designed a streamlined flow that allows rapid comparison of complex text outputs without cognitive overload.

FreeSEED entry screen: the woven knot logo above the tagline 'Left or right? You decide. Help FreeSEED grow and make AI understand Taiwan better,' with a primary 'Walk me through the Arena' button above a secondary 'No, I'll try the Arena myself' button

Both entry buttons name the path rather than the destination — a guided walkthrough, or a route straight into a first comparison — so the choice is legible at the moment it is offered.

Blind comparison screen showing two AI responses side by side to a question about curry and rice, with four voting buttons below

The blind-test comparison — two model responses evaluated side by side before any ritual framing is introduced.

Designing to the Label Schema

The vote resolves to four outcomes — Response 1 is better, Response 2 is better, Tie, and Neither is good. That taxonomy wasn't a UI decision: the model developer set it based on what the training pipeline could actually consume, and my job was to design against it as a fixed constraint rather than propose my own categories. This is the part of the work that doesn't show up in a screen — every layer built after this point, from the blind comparison itself to the ritual framing below, had to confirm one of those four judgments, not replace or blur them.

The Zhijiao System – Digital Folklore

Instead of the generic Western binary of "Thumbs Up/Down," we introduced a vernacular design metaphor based on Taiwanese folk religion: Zhijiao (Moon Blocks).

Sheng Bei (聖茶 - Divine Answer) is used for "Good/Accepted" — it implies the model gave a blessing or a correct, harmonious answer. Xiao Bei (笑茶 - Laughing Answer) is used for "Bad/Hallucination" — it implies the model is "joking" or speaking nonsense.

This design choice was intended to reduce the friction of the task — turning the sterile act of "data labeling" into a familiar, ritualistic action that respects the user's cultural context while still letting us label what users meant.

The evidence behind it is thinner than I would like. The metaphor drew warm reactions from early users, older participants especially, but that was unstructured feedback gathered while demoing — no participant count, no control. So the ritual is a hypothesis with anecdotal support, not a result. The study that would settle it is specific: ritual framing against a plain thumbs-up/down control, measuring comprehension and label quality, and I would run it before committing further design effort to the metaphor.

Post-vote confirmation panel with a split-circle moon-block glyph, the text 'Your judgment was recorded,' the chosen response, a training explanation, and an illustrative community snapshot with vote-distribution bars

The post-vote ritual layer — a moon-block glyph confirms the vote, followed by a plain-language explanation of what the judgment trains and a community snapshot revealed only after voting, to avoid conformity bias.

Information Architecture

The Trinity of Interaction: to sustain high-quality human feedback without losing users' interest, the architecture is divided into three pillars — Chat Mode for blind evaluation using ritual metaphors, Social Mode for community ranking of cultural trends, and Mission Mode for targeted stress-testing. Only the first two were built; Mission Mode stayed a specification.

Information architecture diagram: a story onboarding page leads into the Home (conversation) screen, which branches left to the Island Map and right to Settings, and downward into left navigation, home navigation, content conversation, and topic display — each broken out into its component parts

The information architecture I designed in Figma. The story onboarding page feeds the conversation home, which anchors the evaluation loop — text input, reply, rating, and sharing blocks — while the Island Map carries the social layer of topic and contributor rankings, and Settings holds account, privacy, and contact. Mapping every screen this way is what kept the ritual framing from quietly adding a step the training pipeline couldn't consume.

Outcome

The Arena interface replaces the sterile "A/B Testing" usually found in engineering with a culturally resonant ritual. We mapped the complex task of RLHF to Zhijiao (Moon Blocks), where Sheng Bei — Divine Approval — stands in for "Good Response." The goal was to lower the cognitive load for elderly and local users, turning the tedious work of data labeling into a familiar, intuitive act of "seeking truth."

Social Mode was designed as a community amplifier: crowdsourced rankings and one-tap sharing that carry a private model interaction into public discourse, so the community rather than the team defines what counts as a "good" answer for Taiwanese society.

Share Your Result modal, previewing a branded card with the question, the chosen verdict beside a moon-block glyph, and the illustrative community snapshot, above a Download Image button

The share preview — a private evaluation becomes a shareable artifact, the mechanism behind Social Mode's community amplification. The card carries the verdict and the community split rather than the raw model responses, so what leaves the app is a judgment rather than an unattributed AI answer.

Mobile view at 390 by 844 pixels of the post-vote ritual panel and community snapshot, fitting without requiring a manual scroll

The same post-vote ritual on mobile — the confirmation, explanation, and community snapshot all land within the viewport without extra scrolling.

Status:The Arena shipped as a working product before the program was discontinued. The full Figma design system remains, and I'm currently rebuilding the interface as an independent build project.

What We Didn't Build, and Why

No live community aggregation. The rebuild has no backend, so every vote distribution on screen is static sample data. Rather than imply a live community, the snapshot is labelled as an illustrative prototype figure — a fabricated consensus about Taiwanese values would undercut the entire premise of the product.

No fifth outcome, and no randomized moon block. The obvious move was to replace the four evaluation controls with Sheng Bei / Xiao Bei outright. We didn't, for two reasons. The training pipeline needed the existing four-outcome taxonomy intact, so the ritual had to explain and confirm a judgment rather than become one. And a real moon block is thrown — its result is chance. Animating a random toss over a deliberate human judgment would have made the user's considered choice look arbitrary, so the glyph only ever confirms the outcome the user actually selected.

Community results stay hidden until after voting. Showing a majority result before a user commits invites conformity and contaminates exactly the independent signal the social layer is meant to amplify. The community snapshot is a reward for participating, not an input to the judgment.

Reflection

The hardest part wasn't the ritual metaphor — it was aligning the chat flow and the interaction design so the feedback data actually came out clean enough to refine the model. Every design choice about how users compared and voted on responses directly shaped the quality of the labels we could feed back into training, so I had to design the interaction and the data pipeline as one problem, not two.

More broadly, the project pushed me to design around the standard engineering approach to AI training. The technical goal was simply to gather human feedback, but the real design challenge was sustaining engagement while keeping that feedback structured and usable. My working belief coming out of it is that gamification lands better when it taps into existing cultural behavior than when it bolts on a generic point system — mapping 'Sheng Bei' (Divine Approval) onto 'good response' was an attempt to make cultural nuance carry real infrastructure load in a high-tech system.

LinkedIn·yushengchen.uxr.build@gmail.com
© 2026 Yusheng Chen. All Rights Reserved.