MankitSze Magazine · Chinese Edition

From Fortune-Telling App to AI System

The product, engineering and leadership decisions behind 觀音AI .

Written and edited by
Mankit Sze
MankitSze MagazineIssue 001 · Behind the Build

From Fortune-Telling App to AI System

觀音AI did not begin as a new AI product.

Its roots go back to an earlier generation of mobile applications, when the main job of a fortune-stick app was straightforward: let someone draw or enter a fortune-stick number, retrieve the corresponding interpretation, and make traditional content conveniently available on a phone.

That problem had already been solved.

Generative AI created a different opportunity.

The interesting question was not whether AI could generate another interpretation of a fortune stick. The existing content already did that reliably.

The interesting question was:

What happens after someone reads the fortune?

A person asking about a relationship is rarely interested only in the literal interpretation of a poem.

They may be trying to decide whether to send another message, whether to wait, whether to leave, whether they are overreacting, or whether they are seeing the situation clearly.

Someone asking about career fortune may really be struggling with whether to stay in a job, confront a manager, accept an offer, or tolerate uncertainty a little longer.

The fortune stick can introduce an unexpected perspective.

But the person still has to understand what that perspective means in the context of their actual situation.

That became the space for AI.


Giving the Fortune Stick and AI Different Responsibilities

One of the earliest product decisions was that AI should not become an electronic fortune teller.

The two systems should have different responsibilities.

觀音靈簽 provides intuition, symbolism and insight.

AI provides reasoning, objectivity, emotional steadiness and practical next steps.

That distinction changed the build.

If the goal had simply been personalized fortune interpretation, a conventional chatbot wrapped around an LLM would have been enough.

But the product needed to do several different things well.

It needed to understand the fortune-stick context without treating it as factual evidence.

It needed to understand the user's real-world situation.

It needed to distinguish what was known from what the user feared, hoped or assumed.

It needed to recognize emotional context.

It needed to consider possible actions.

And it needed to combine all of those perspectives into a response that felt coherent rather than mechanical.

That is why 觀音AI evolved into a multi-agent system.


Building a Team Instead of a Single Assistant

觀音AI is built around multiple AI agents collaborating on the same problem.

The agents do not exist as characters for the user to select.

They exist because the problem itself contains different responsibilities.

A single user message may simultaneously contain a question about the fortune, a real-world decision, an emotional reaction and an implicit request for advice.

Rather than expecting one undifferentiated assistant to perform every kind of reasoning equally well, the system separates responsibilities.

One perspective can concentrate on what the fortune contributes.

Another can examine the situation rationally.

Another can consider emotional context.

Another can concentrate on what the user can realistically do next.

The important part is not the number of agents.

It is that different responsibilities are made explicit.

This makes 觀音AI less like a chatbot with an elaborate personality prompt and more like a small AI team working on behalf of the user.


Collaboration Matters More Than Specialization

Splitting responsibilities between agents is only useful if those agents actually collaborate.

If five agents independently produce five answers and the application simply joins them together, the system has gained complexity without gaining much intelligence.

The different perspectives need to affect one another.

Consider a user saying:

My girlfriend said she needs some space.

It's been almost two days since she replied.

Should I send another message, or should I leave her alone?

The fortune-stick perspective may introduce an idea such as patience, restraint or waiting for the appropriate moment.

That can be useful.

But it cannot become evidence that waiting will definitely produce a good outcome.

A more rational perspective can point out that two days of silence does not reveal what the girlfriend is actually thinking.

An emotionally aware perspective can recognize that the urge to send another message may be driven partly by the discomfort of uncertainty.

A practical perspective can then ask whether another message would genuinely improve communication or simply reduce the user's anxiety for a few minutes.

These perspectives are not independent answers.

They constrain and improve one another.

The symbolic perspective can introduce an idea.

Rational analysis can prevent that idea from becoming an unsupported conclusion.

Emotional understanding can explain why one option feels urgently necessary.

Practical reasoning can turn the combined understanding into something the user can actually do.

The final response should feel like one thoughtful answer rather than the transcript of an internal meeting.


One 觀音AI on the Outside

The multi-agent architecture is deliberately hidden behind a simple conversational experience.

The user does not have to understand agents.

They do not have to decide which specialist should handle a relationship question.

They do not have to switch between a fortune agent, reasoning agent or emotional-support agent.

They talk to 觀音AI .

Internally, the system can be specialized.

Externally, it remains one product.

That became an important design principle:

Internal complexity is worthwhile when it creates external simplicity.

Multi-agent collaboration is therefore an implementation strategy, not an additional interface for the user to manage.


Why Not Use One Giant Prompt?

A sufficiently capable language model can be given a long list of instructions.

It could be told to interpret the fortune, remain objective, recognize emotions, challenge assumptions, suggest actions and produce an appropriate final response.

That is certainly easier to implement.

But it also concentrates many different responsibilities inside one opaque reasoning process.

When the result is poor, it becomes difficult to understand why.

Did the system misunderstand the fortune?

Did it accept the user's assumption as fact?

Did it overlook the emotional context?

Did it give sensible analysis but an unrealistic recommendation?

Did one instruction receive less attention because too many other instructions were competing with it?

Separating responsibilities gives the system structure.

Each responsibility can be reasoned about independently while still contributing to the same outcome.

The architecture can also evolve.

A responsibility that consistently adds little value can be changed or removed.

A weak capability can be improved without redesigning the entire product.

Different responsibilities can potentially use different models or techniques.

The architecture is therefore not based on the assumption that one model must be equally good at everything.


The System Instruction Became a Product Artifact

Another important part of the build was recognizing that the system instruction was not merely configuration.

It encoded product decisions.

What should 觀音AI do when the fortune appears to contradict the user's expectations?

How should it distinguish symbolic interpretation from factual reasoning?

When should it challenge the user's assumptions?

When should it provide emotional support before suggesting an action?

How strongly should it recommend a next step?

What should it refuse to claim?

How should the different AI responsibilities work together?

These are product questions expressed through instructions.

The quality of an AI product therefore depends not only on code and model selection, but also on how clearly the product's responsibilities are defined.

Prompt design became part of system design.


The User Prompt Is More Than a Question

The user's latest message is only one part of what the system may need to understand.

A useful response can depend on the relevant fortune-stick context and what has already happened in the conversation.

This becomes especially important over time.

Suppose someone previously discussed whether to leave a job.

The system helped them think through the situation and suggested speaking with their manager before making a decision.

When they return several days later, starting again from:

How can I help you today?

throws away much of the value created by the earlier conversation.

A better system can continue the process.

Were you able to speak with your manager?

Or:

Has anything changed since you were considering leaving?

The conversation becomes cumulative rather than disposable.


Starter Prompts Became Part of the Intelligence

This also changed the purpose of starter prompts.

In many AI products, starter prompts are examples designed to help users understand what they can ask.

That makes sense for a first conversation.

It becomes less useful after the system already has context.

For returning conversations, useful prompts can instead ask for new developments.

They can ask whether the user followed through on a previous action.

They can suggest another practical step.

They can offer emotional support if the situation has not improved.

They can even change direction when the previous approach was not useful.

The prompt is no longer merely:

Here is something you could ask the AI.

It becomes:

Here is a useful place to continue the work we were already doing.

That makes conversation history part of the product rather than merely an archive of old messages.


Useful Memory Creates Sensitive Memory

The more useful the system becomes at remembering context, the more sensitive that context can become.

A fortune-stick number is relatively impersonal.

A conversation about why someone believes their relationship is collapsing is not.

Neither is a discussion about career failure, money, family conflict, regret, fear or something the user has never told another person.

This created an architectural consequence.

Privacy could not remain a policy written after the product was finished.

It had to become part of the build.


Privacy Changes the Quality of the Conversation

Conversational AI depends on context.

If users censor themselves, the system reasons from incomplete information.

The subjects where 觀音AI may be most useful are also subjects people may be most reluctant to send somewhere else.

That led to an important realization:

Privacy does not only protect the conversation after it happens. It can affect what the user is willing to say in the first place.

If someone believes their private thoughts are being uploaded, stored or analyzed elsewhere, they may leave out exactly the details that matter.

If they feel able to speak honestly, the system receives better context.

Better context makes better reasoning possible.

Privacy therefore became part of product quality.


Moving the AI Onto the Device

The strongest architectural response was to avoid sending the private conversation away simply to obtain an AI response.

That meant running AI on the user's device.

The conceptual benefit is straightforward.

The user's conversation can remain on the phone.

Conversation history can remain on the phone.

Inference can happen on the phone.

The architecture itself can support the privacy promise.

But the engineering consequences are substantial.

A cloud AI service can use infrastructure specifically designed for large models.

An iPhone has finite memory, finite storage, finite computational resources and a battery.

The AI system has to live within those constraints.


Model Selection Became Product Engineering

Once inference moved onto the device, choosing a model was no longer simply a question of which model produced the best answers.

A larger model may reason better.

It also consumes more storage.

It may require more memory.

It may respond more slowly.

It may exclude older devices.

It may create larger runtime caches.

A smaller model may be faster and easier to deploy but fail to provide sufficiently nuanced reasoning.

The relevant question therefore became:

What level of intelligence is useful enough for the product within the physical constraints of the devices users actually own?

That is a different optimization problem from choosing a cloud model.

There is no single dimension to maximize.

Reasoning quality, model size, memory, latency, device compatibility and privacy all interact.


Multi-Agent Reasoning Has a Cost Too

The same trade-off applies to the agent architecture.

More agents are not automatically better.

Every additional reasoning step has a cost.

On a cloud service, additional inference may primarily become a financial and latency problem.

On-device, it can also become a compute, memory, thermal and battery problem.

That means multi-agent architecture needs discipline.

An agent should exist because separating that responsibility improves the final result enough to justify the additional complexity.

The objective is not to build the largest possible AI organization inside the phone.

The objective is to find the smallest collaboration structure that produces meaningfully better reasoning.

This is especially important on mobile hardware.

Architecture has to respect physics.


AI Was Also Changing How I Built the Product

There was another AI transformation happening at the same time.

AI was becoming part of the development workflow itself.

Instead of manually translating every product decision directly into implementation, I increasingly worked through explicit artifacts.

Product thinking could be captured in Markdown.

Architecture decisions could be documented.

Screenshots could become visual references.

Screenshot specifications could describe intended output.

Those artifacts could then provide context to coding agents working against the actual codebase.

This changed where much of the engineering effort went.

Writing every line of code personally became less important.

Defining precisely what should be built became more important.


The Bottleneck Moved From Code to Intent

AI can generate implementation quickly.

It can also generate the wrong implementation quickly.

The quality of the result depends heavily on the quality of the context surrounding the task.

What is the product trying to accomplish?

Which decisions have already been made?

Which ideas are only experiments?

What must remain unchanged?

Which artifact is the source of truth?

What are the architectural constraints?

What does success look like?

When implementation becomes cheaper, ambiguity becomes more expensive.

This changed my role in the build.

More time moved toward defining the problem, establishing constraints, reviewing results and maintaining coherence across the system.

AI accelerated implementation.

It did not replace technical judgment.


Generated Ideas Are Not Shipped Features

AI-assisted development also creates a new form of project-management risk.

Generation is cheap.

A brainstorming session can produce ten product ideas.

An image model can produce several visual directions.

A coding agent can implement an experimental feature.

A language model can write a detailed architecture that sounds entirely plausible.

None of those things necessarily represents the actual product.

This makes source-of-truth discipline increasingly important.

An idea that was discussed is not the same as a decision.

A generated screenshot is not the same as an approved design.

A proposed architecture is not the same as the architecture that was implemented.

A feature mentioned in a document is not necessarily a feature that shipped.

The faster AI produces possibilities, the more important it becomes to distinguish exploration from reality.


The Requirements Emerged Through the Build

觀音AI did not follow a straight line from specification to implementation.

Each decision exposed another problem.

The existing fortune-stick experience suggested conversation.

Conversation created a need for richer reasoning.

Richer reasoning led to separating responsibilities across collaborating agents.

Ongoing conversation created a need for continuity.

Continuity created a need for memory.

Memory made privacy more important.

Privacy pushed AI inference onto the device.

On-device inference introduced model, memory, storage and performance constraints.

Multi-agent reasoning introduced its own compute and orchestration costs.

Those constraints fed back into product decisions.

The architecture did not merely implement a fixed specification.

Building the architecture helped reveal what the product needed to become.


What I Chose Not to Build

Some of the most important engineering decisions were decisions not to use AI.

I did not replace deterministic fortune-stick retrieval with generation.

The existing mechanism was already better at that job.

I did not turn AI into a machine that pretends to know the future.

I did not treat the fortune as evidence that a predicted outcome will occur.

I did not expose internal agents as personalities that users need to manage.

I did not add agents simply because multi-agent systems were technically possible.

I did not treat conversation history merely as a mechanism for increasing engagement.

And I did not treat privacy as something that could be solved entirely through policy language.

Each of these constraints removed possibilities.

That was useful.

AI makes adding things easy.

Good product engineering still requires knowing what not to add.


The Architecture Followed the Product

觀音AI eventually became a system involving traditional fortune-stick content, conversation context, system instructions, user prompts, multiple collaborating AI agents, response synthesis, conversation continuity, local storage and on-device inference.

But none of those technologies explains why the system exists.

The architecture began with a much simpler decision:

觀音靈簽 and AI should do different jobs.

The fortune stick provides intuition, symbolism and insight.

AI provides reasoning, objectivity, emotional steadiness and practical next steps.

Once those responsibilities became clear, the rest of the system started to take shape.

Different kinds of reasoning could be separated into collaborating agents.

The agents could challenge and complement one another.

Their work could be synthesized into one response.

Conversation could continue over time.

Privacy became essential because those conversations could become deeply personal.

On-device inference became valuable because privacy was part of the product rather than an afterthought.

The architecture followed those decisions.

Not the other way around.


Building AI Without Making AI the Product

The easiest way to build an AI product today is to begin with AI.

Choose a model.

Choose an agent framework.

Choose an inference architecture.

Then find something for the technology to do.

觀音AI evolved in the opposite direction.

The original product already did something useful.

The first question was what remained unsolved.

That led from static interpretation to conversation.

Conversation exposed different reasoning responsibilities.

Those responsibilities led to multi-agent collaboration.

The sensitivity of the resulting conversations made privacy fundamental.

Privacy moved the intelligence onto the device.

And the limitations of the device forced every AI decision back into the realities of software engineering.

Internally, the system became considerably more sophisticated than the fortune-stick application from which it evolved.

Externally, the goal remained simple.

The user should not have to understand the model.

They should not have to understand the agents.

They should not have to understand orchestration.

They should not have to understand the architecture.

They should simply have a place where they can say what is really happening, receive more than one kind of perspective, and leave with a clearer idea of what they might do next.

That is the part of the build that matters.

MS

End of Behind the Build