Building Aveline: A Multi-Agent Concierge Platform

Building Aveline: A Multi-Agent Concierge Platform

October 3, 2026 30 mins read

Aveline is a multi-agent commerce and concierge platform for semi-luxury boutiques in Sri Lanka. A customer messages the boutique on WhatsApp. A floor associate works a Flutter app on the shop floor. The owner works from a React dashboard, watching the approval queue.

Between those three people sits the part I actually want to write about: an ASP.NET Core modular monolith, a Python agent service running three LangGraph sub-graphs, and a PostgreSQL database that holds both the business data and the vectors the agents reason over.

This post is not a tour of the final architecture. It is the record of the decisions that held up, and the ones that did not — including two bugs that were invisible in the design document and obvious the moment a real message arrived.

Context

Aveline was built for the SE3090 group assignment with two teammates, split into three graded vertical slices: Customer Concierge & Memory, Visual Intelligence & Sourcing, and Commerce Validation & Optimization. I owned Slice 1 — the customer relationships, memory and messaging side — but the seams between the slices are where the interesting failures live, so this post covers the whole system where it has to hang together.

The problem: a boutique cannot answer everyone

A semi-luxury boutique in Colombo has a specific failure mode. The owner knows her regulars by name, remembers that one client will not wear nylon, and knows which piece is sitting on the rail too long. What she cannot do is hold all of that while twenty WhatsApp threads are open, price a basket correctly against a 25% margin floor, and remember that the same customer asked about a Kanjeevaram saree four months ago.

That is the job we gave Aveline. Concretely, a single inbound message has to be able to trigger:

  • identifying the customer from a phone number, with consent state respected;
  • retrieving what the boutique already knows about her, by meaning rather than keyword match;
  • searching the catalogue for what she is asking about, including from a photo;
  • pricing the resulting basket against the boutique’s own business rules;
  • and pausing for the owner when the rules say the decision is not the agent’s to make.

That is a workflow, not a prompt. Which is where the first architectural decision came from.

Modular monolith, not microservices

The obvious instinct for a system with three distinct domains and a Python reasoning service is microservices. We wrote it down as an option and rejected it.

One Aveline.Api project, with a folder per vertical slice. Each slice owns its own entities, its own EF Core configurations, its own services, and its own endpoint group. Nothing stopped one slice from reaching into another’s tables except convention, and that was a deliberate trade: for a small team on a fixed timeline, one deployable, one migration history and one local dev story were worth more than enforced isolation.

The real boundary that we did enforce was the one between the two runtimes. The Python agent service never writes to the database and never calls a third party directly. Every piece of business logic — pricing, consent, persistence, permissions — lives in the .NET API behind internal endpoints guarded by an X-Internal-Token header. The agent reasons; the API decides.

That line is what made the rest of this possible, and it is the line the commerce agent’s approval rule is built on. A language model should not be the authority on what a valid order is.

Three agents, one workflow

The agent service runs a single top-level LangGraph graph that routes to three specialist sub-graphs. The specialists are named for the people they stand in for: Ava for customer memory, Elle for visual intelligence, Lina for commerce.

Two things to notice. First, there is no bind_tools anywhere in this codebase. Every “tool” is a deterministic node step calling a ToolRegistry; the model never decides to call something, and what each agent can call is fixed by what its module imports. Second, with agent_llm_enabled=false or no API key configured, the whole thing still runs on a rule-based path. That guarantee is not a nicety — it is what keeps CI green without a model in the loop, and it is what let us develop offline.

Both of those properties caused a bug. The second one, in particular, caused the worst one.

The bug that vetoed every new customer

Here is a real exchange from the integration, verbatim:

Customer: Hello there. Are there any pinkish gowns in your collection? Aveline: I couldn’t find a customer with that name. Could you share their phone number so I can look them up?

The customer was messaging from their own WhatsApp number. The number was right there in the request payload. Aveline was asking them to provide it.

The diagnosis took a while because the symptom looked like a customer-resolution problem, and it was not. Customer resolution worked fine — an existing customer’s number resolved correctly. The problem was what we did with the answer. The original pipeline had a hard gate: any ambiguous or not_found customer resolution routed straight to formulate_response, and formulate_response discarded every specialist output when that happened. The only thing that could ever be rendered was a clarification block.

For an inbound WhatsApp message, “I don’t know this customer” is the expected state for every first-time contact. We had made the single most common case a veto on the entire workflow. No product question from a new customer could ever be answered.

There was a second layer underneath it. Routing was a keyword table, and the table listed dress, saree, blouse, outfit, party, bluish, size, stock, item, photo, image, picture, matching. Gowns were not in it. pinkish was not in it, though bluish was. classify_by_rules returned general_inquiry, routing was ["memory"], and the visual agent never ran.

We could have added gown. We had already written a hybrid intent gate — deterministic rules first, an LLM classifier for ambiguous input — and the classifier had no call site. Adding a synonym would have fixed the instance and left the class intact, one synonym away from failing again.

What we replaced it with

Routing authority moved from the static table to a supervisor node that runs once per turn, sees a bounded conversation window plus customer state, and emits a structured plan. The supervisor decides; a deterministic driver executes. LangGraph’s edges stay fixed and the driver iterates the agent list the supervisor returned, so routing becomes data rather than a graph rewrite — which keeps the workflow auditable and, critically, keeps it resumable under the existing checkpointer.

The rules are now a cheap pre-filter for genuinely unambiguous input and the complete fallback when there is no model — not the authority. The pre-filter is deliberately tiny, and there is a scar behind that: the first implementation listed the three intents the model was consulted for, which meant a keyword hit was treated as proof the message was unambiguous. _RULE_KEYWORDS matches "prefers" anywhere, so a staff instruction to record a preference — “please add a note for this customer: he prefers green tea” — routed as a customer stating one, the supervisor never ran, and the extractor then found nothing first-person to store. The pre-filter is now the narrow enumerated set it should have been from the start.

The other half of the fix was conceptual: customer resolution stopped being a gate and became a fact. resolve_customer now records resolved / ambiguous / not_found / no_signal; whether that fact requires asking the customer something is the supervisor’s decision. A run can produce a clarification without discarding specialist output it already produced. And a not_found derived from the sender’s own context must never ask for the sender’s number.

That last sentence is a one-line invariant in the ADR. It took a production bug to earn it.

Context is layered, and deletion is a UX decision

Handing the supervisor conversation history sounds trivial until you do it wrong. The context model settled into four layers with different retention policies:

LayerContentsRetention
Recent turnslast N turns, verbatimbounded by a token budget
Thread summaryrolling narrative of what left the windowreplaced on compaction
Durable factspreferences, events, constraintspermanent, in pgvector
Live dataprofile, stock, pricingnever retained; fetched on demand

The rule that matters most: turn-pair integrity. Naive FIFO eviction separates a question from its answer. Drop “which dress did you mean?” while keeping “yes, that one” and the referent is destroyed — the assistant is now confidently wrong rather than merely uninformed. So the window trims on turn boundaries, never inside a pair, and a small set of pinned slots (the item under discussion, a stated budget, an event date) survives trimming while the thread depends on them.

And a point that the design doc got wrong the first time: budget is measured in tokens, never in message count. With three specialists in the graph, the dominant consumers are tool results — catalogue searches, image analyses, margin computations — not chat turns. A tidy 10-message window will not compensate for a verbatim 4,000-token inventory dump coming back from a search every turn.

Memory that survives a conversation

Ava’s job is to remember the small things. “Nothing nylon, or spandex, I don’t like them, they make my skin itchy” is exactly the kind of sentence a boutique owner would remember and a database would lose.

Memory lives in PostgreSQL 16 with pgvector. Embeddings are 1536-dimensional (text-embedding-3-small, behind an IEmbeddingService seam) in a vector(1536) column with an HNSW cosine index.

One decision there is worth stealing: the vector column and its index are not part of the EF Core model. They are created by the migration via raw SQL — vector(1536) plus an HNSW vector_cosine_ops index — and read through raw SQL in the repository. The alternative, mapping a CLR vector type onto the entity, breaks EF Core’s in-memory provider, which cannot map vector, and that would have broken roughly four hundred unit tests for the sake of a nicer model. Real-Postgres behaviour is verified separately with Testcontainers against pgvector/pgvector:pg16.

Two requirements, two test harnesses

The trade is deliberate: the in-memory provider covers everything that does not touch the vector column, and a Testcontainers fixture covers the raw SQL that does. If you have a Postgres-specific feature and a large in-memory suite, keeping the feature outside the ORM model is often cheaper than fighting the provider.

Search itself is hybrid. A cosine leg and a full-text leg run as one statement over the same tenant, customer, live and unexpired filter, and their results are fused with Reciprocal Rank Fusion:

The reason is not that hybrid is fashionable. It is that the two legs answer different question shapes, and memory queries are overwhelmingly exact tokens — a garment name a customer mentioned (“Kanjeevaram”, “Banarasi silk”), a label-like phrase, a term a paraphrase-trained embedding ranks mid-pack. A dense-only leg blurs a specific garment name into a general “elegant clothing” region, and the boutique experiences that as it forgot the saree I told it about. RRF is rank-based because cosine similarity and ts_rank_cd are not on a comparable scale, so a hit found by only one leg still has to place.

The details that took the longest to get right were all about degradation:

  • If the embedding provider is unavailable, the lexical leg answers alone and VectorRank comes back NULL — the degradation is visible on the wire rather than silent.
  • minSimilarity gates the dense leg only. A floor applied after fusion would delete exactly the lexical-only hits the hybrid exists to surface.
  • MemorySearchRequest.Mode accepts hybrid / lexical / vector, and an unknown mode returns 400 rather than defaulting. A misspelled single-leg request silently answered by fusion would be read as a leg measurement.

Writing facts into memory has the same shape. A deterministic parser recognises three shapes in the current message — a first-person preference, a dated event, a quoted complaint — and a model-driven extract node runs between retrieve and persist to catch what the regexes cannot. The extraction node owns no writes; persist owns every write. If the model is missing, fails, or returns something unparseable, the run loses facts and never loses the run.

The category vocabulary is closed — preference, event, observation, complaint, constraint — because a model free to name its own category invents a taxonomy one row at a time, and that table is read by people. And the transcript resolves references; it is not a source of facts. Each stored fact is one self-contained sentence that starts with the customer’s name, so it makes sense with no context beside it. A model handed a conversation will summarise it, and a summarised conversation is not a fact about this customer.

Human in the loop: the pause is the whole node

The rule that defines the whole platform: when a basket breaches a business rule, the agent stops and a human decides. The defaults live in the API, not the agent.

The breach returns pending_approval with needs_approval, approval_type, approval_reason and the triggered rules, and no payment link. The API creates the Order only at that moment — status pending_approval — and then the approval queue entry that references it, carrying the thread id and conversation id. Creating the order only at the pause keeps its existence tied to a decision that actually needs one, instead of writing speculative orders for every enquiry.

The part I would tell anyone building this to get right is what happens on the way back.

The original resume used a re-query: the API posted the decision to /agents/query with text like "[Human Approval Decision: approve]". That fails in three separate ways at once. It re-runs the entire pipeline, so it re-sends no line items and evaluate_deal returns skipped before it ever reads the decision. The vocabularies do not match — the API sends approve/reject/revise, the graph matches approved/rejected. And the re-query text carries no commerce signal, so the supervisor routes to memory only and the commerce agent never runs.

It also is not a resume. It is a re-run, and every side effect before the pause executes again unless each one is separately guarded for idempotency.

So resume is now a real checkpoint resume:

Four things here cost us real time and are worth writing down.

The pause is a node of its own. LangGraph re-executes the node that called interrupt() when the graph resumes. Putting the pause inside run_commerce_agent would re-evaluate the deal — and its tool calls — on every decision. Making the pause the entire node means the evaluation before it is checkpointed and never re-run. That is a one-line structural decision with a correctness consequence.

aget_state reads the checkpointer off the compiled graph. Passing the saver to ainvoke works for running the graph but not for asking whether the run is waiting, so the checkpointer has to be compiled into the graph.

The decision vocabulary is one published set, shared by the API and the graph, with a single translation point on the .NET side. An approve vs approved mismatch is a class of bug that only appears at the seam between two runtimes, and it appears as silence.

The resume must be idempotent, and it is enforced twice — a service check for the common retry, and a partial unique index on (OrganizationId, ThreadId) WHERE Status = 'pending' for the race, scoped to pending rows so a thread can legitimately place another order once the first is decided. The migration scaffolder wanted to backfill legacy rows with an empty string, which would have made every one of them collide on the new index; they were set to legacy-{orderId} instead.

And one more, because it is the kind of thing that only exists at the seam between an agent and a database: a staff order has no agent run behind it. An order created from the dashboard gets a generated thread id and a null conversation id, and that null is the marker. A decision on such a row is complete without a resume, and the code skips rather than attempting to resume a checkpoint that was never written.

Test-first, because the seams are invisible from the design

The lesson above about testing seams did not arrive at the end. It was written into the project’s rules on day one.

.agents/rules/Rules.md is the standing instruction file every contributor — human or agent — works against, and section 7 is not a suggestion:

Rules.md § 7, Test-Driven Development
  1. Define the expected behavior.
  2. Write the test.
  3. Run the test and confirm that it fails for the expected reason.
  4. Implement the minimum code necessary to satisfy the test. 5. Run the test again. 6. Refactor while keeping the tests passing. — Do not write implementation first and create superficial tests afterward.

Step 3 is the one people skip, and it is the one that matters. A test you never watched fail is a test that might be asserting nothing — passing because the code was already right, or because the assertion is vacuous. “Confirm it fails for the expected reason” is the difference between a test suite and a pile of green checkmarks.

It worked, and not in the way I expected it to. The clearest example came from the mobile app’s settings screen:

The picker’s Cancel sat below the visible viewport: the sheet overflowed by 182px under the test font. The fix was a real one rather than a test tweak — the options (and the detail fields) now scroll while the actions stay pinned, so a sheet can never hide the way out.

Nobody filed that bug. There was no device in the loop and no screenshot. A test written before the widget existed, rendered at a font size the design had not accounted for, reported a 182-pixel overflow — and the temptation in that moment is to widen the tolerance or delete the assertion. We changed the layout instead: the way out of a modal can never scroll off it. That is a design invariant now, and it was discovered by a test, not by a user.

The same log records a second finding worth keeping honest about: the tests did not catch everything. A device review of the running screen found Store role: org:boutique_owner rendered twice — an internal claim id printed as a value. No widget test failed, because every widget test agreed with the code. It took looking at the screen.

That is the honest shape of the payoff. Test-first is not a guarantee; it is a way of moving the moment of discovery earlier — earlier than the device, earlier than the demo, earlier than the customer. It cannot see what the test author did not think to assert.

The gates, and the one thing they cannot do

Four independent codebases means four test harnesses, each with a floor enforced in CI:

StackToolingGate
Backend APIxUnit, Moq, WebApplicationFactory, Testcontainersline coverage ≥ 30%
Agent servicepytest, respx, fakerediscoverage ≥ 90%
Web dashboardVitest (@vitest/coverage-v8)lines 80 / functions 70 / branches 70 / statements 80
Mobile appflutter_testline-coverage floor at 81%
Shared contractone JSON contract, asserted from both C# and Python—
End to endPlaywright against a composed stack—

The asymmetry in those numbers is deliberate rather than sloppy. The agent service holds the highest floor because it is the least deterministic thing in the system — a graph that routes, retrieves and composes is exactly where an untested branch becomes a wrong answer. The .NET API holds the lowest, because a 30% line floor across a controller-and-repository codebase catches the untested path, while the integration suites that actually matter run against a real Postgres through Testcontainers. A single global number would have been easier to state and useless to enforce.

And the gate that I would carry into any team: no build artifact is produced unless every test stage passed. Not a warning, not an allow-failure, not a dashboard someone checks. The CI spec states it with no bypass conditions, and that is the whole point — a quality gate with an escape hatch is a quality gate that gets used.

The gate that no coverage number can express is the one I wrote about above: the seams. A 30% line floor on the API and a 90% floor on the agent service will both pass while the contract between them is wrong. That is why tests/contracts/ exists — one JSON contract asserted from both sides — and why the scenario script that replays a real customer message end to end through the composed stack is the thing I run first.

Deployment: deciding the topology before writing the template

The assignment required the whole cross-platform workflow — Flutter → API → agent → database → React → back — to run against a hosted backend, inside a < $100 budget, by three full-time students. We wrote that down as an ADR before choosing anything, which is the only reason the choice was defensible later.

Microsoft Azure, and specifically the free tiers: Container Apps, PostgreSQL Flexible Server Burstable B1ms, Key Vault, Log Analytics. Railway and Render were rejected because neither has a free tier that carries this workload. The tempting option — Vercel + Supabase + best-of-breed everything — was rejected for a reason I would repeat: it splits the stack across vendors and weakens the one thing the brief actually asked us to demonstrate, a single coherent deployment pipeline.

That decision held for the API and the agent. The web dashboard is on Vercel anyway, and the reason is worth recording because it is the kind of thing the plan cannot anticipate: Azure Static Web Apps could not be created in any permitted region. The student subscription exposed five regions — indonesiacentral, indiasouthcentral, koreacentral, malaysiawest, uaenorth — and the whole stack had to live in one. So frontend/web/vercel.json exists in the repository, the dashboard is at aveline.gravora.dev, and ADR-006 still claims Static Web Apps because it was never amended. Documentation catching up to reality is the theme of this build.

The live topology, as it actually runs:

ResourceNameDetail
Resource grouprg-avelinemalaysiawest — the only permitted region hosting the whole stack
Container Apps envcae-avelineConsumption-only, no VNet
Container Appaveline-apiexternal ingress on 8080, minReplicas: 1
Container Appaveline-agentinternal ingress only on 8000, minReplicas: 1
PostgreSQL Flexible Serverpg-avelinev16, Burstable B1ms, 32 GB, vector + btree_gist
Azure Managed Redisredis-avelineBalanced B0, NoCluster, NoEviction, TLS on :10000
Container RegistryavelineacrBasic, adminUserEnabled: false
Key Vaultkv-aveline-mwRBAC authorization
Log Analyticslog-aveline30-day retention

Three decisions in that table are load-bearing, and all three contradict the original plan.

The agent has internal ingress only. It is not reachable from the internet at all — the only thing that can call it is the API, through Container Apps’ internal DNS, carrying the same X-Internal-Token the local setup uses. That is the two-runtime boundary from earlier in this post, enforced by the platform rather than by convention.

Nothing scales to zero, and the reason is in a comment in the Bicep. The API is not a stateless request handler: it runs billing rollups, retention sweeps and metric collectors on in-process timers, and it holds the Redis PSUBSCRIBE subscriptions for the event bus. Scaled to zero, the timers stop ticking and any event the agent publishes while the API sleeps is simply lost. The original ADR had minReplicas = 0 for cost; the template now pins minReplicas: 1, maxReplicas: 1 on both apps, with the reasoning written directly above it:

Note what that is: a scaling parameter that is really a design decision, because the application was built with in-process timers and a pub/sub subscriber that both assume a warm process. Horizontal scaling would break it the same way.

Redis is NoEviction, not the default. Locally we run redis:7-alpine with its default VolatileLRU policy, which is fine until you remember what is stored with a TTL: job locks, idempotency keys, and the approval-resume guards. An eviction policy that can drop a lock to reclaim memory is an eviction policy that can turn an idempotency guarantee into a duplicate order. Production pins NoEviction and treats a full Redis as an error to be fixed rather than a cache miss to be absorbed.

Shipping it

The deploy workflow is deliberately separate from CI. CI validates every push and pull request and must stay fast and secret-free; the deploy workflow ships, so it runs only on workflow_dispatch, and only from a ref whose OIDC subject is trusted.

Authentication uses a user-assigned managed identity with a GitHub OIDC federated credential — no client secret exists at any point. That was forced by the tenant: allowedToCreateApps is false, so az ad sp create-for-rbac fails with a 401, and app registration is unavailable to ordinary users. A managed identity is created by the resource provider rather than the Graph API, so it is unaffected. The federated credential pins an immutable subject per branch, which means a workflow triggered from an untrusted ref cannot log in even if someone edits the YAML.

Migrations are a separate, opt-in job rather than something the app does at startup. The app only auto-migrates when IsDevelopment(), so production needs an explicit step, and that step refuses to run on an empty connection string:

The roll itself updates the agent before the API — so the callee is never older than the caller — and then verifies rather than assumes: it polls https://<fqdn>/health up to thirty times at twenty-second intervals until the app reports healthy, and asserts that the running image tag is the one that was just pushed. A roll that silently left the old revision serving traffic would otherwise look like a success.

Secrets never enter the image or the repository. Sixteen Key Vault secrets are referenced by Container Apps secretRef entries resolved through the app’s system-assigned identity, and the Bicep template itself reads those values from the deploying shell at deploy time:

What is not deployed, and why I am saying so

Three gaps, stated plainly because a deployment section that only describes what works is marketing.

The observability stack is local-only. Prometheus, Grafana, Jaeger, the OTel collector and the Postgres exporter all run in docker-compose.yml and none of them exist in the Bicep. Azure Application Insights would be the production answer; it is deferred. The application is instrumented either way — the OpenTelemetry wiring does not change — but in production those traces currently go nowhere.

The Bicep has never been deployed to an empty resource group. It compiles with zero warnings and every resource API version was taken from a live ARM capture rather than from memory — but the live infrastructure was created by hand with the az CLI, and the template is a reviewed starting point, not proven infrastructure. Its own README says exactly that.

Everything was deployed by hand first. There was no deploy job in this repository until late in the build. That is the honest order of operations for a project like this, and it is also why the environment drifted from docs/deployment.md: the documentation described a plan, the infrastructure was built by a person, and the two only converged when the pipeline was finally codified.

The one habit I would keep

Migrate in its own gated job, and make it refuse rather than guess. An application that auto-migrates on boot will eventually migrate against the wrong database — and the failure is silent, because the app starts up perfectly happily either way.

What building this taught me

Six weeks, 478 commits, and a system that reads as much smaller than it is. Three lessons that I would carry into the next one.

The prompt is not the architecture. Every one of the failures above was a routing, ownership or lifecycle problem that no amount of prompt engineering would have fixed. The pinkish-gowns bug was a gate. The re-query bug was a vocabulary mismatch. “It forgot the saree” was a dense-only retrieval leg. Agents are easy to prototype and hard to place — the work is deciding what each one may read, what it may write, and who owns the decision when it is wrong. Our answer was firm: the agent reasons, the API decides, and the agent never touches the database.

Determinism is a feature, and a cheap one. A single agent_llm_enabled switch keeps the entire workflow rule-based. That gave us an offline development mode, a CI path with no API key and no flake, and a graceful degradation story where a provider outage costs quality rather than availability. It also forced every LLM call to justify its existence against a working fallback, which is a healthy thing to be forced to do.

Test the seams, because that is where the bugs are. Every coverage floor in the table above was green when both of this post’s bugs shipped. A test suite proves that each side does what its author believed; it says nothing about whether two sides agree. Budget for the tests that cross a boundary — the shared contract, the composed stack, the scenario script — as a separate line item, not as a rounding error on unit coverage.

If you take one thing away

Put the pause in its own graph node, make the resume a real checkpoint resume, and write the decision vocabulary down in one place. Everything else in a human-in-the-loop agent system is negotiable. Those three are not.

Where it goes next

Aveline is functional end to end: WhatsApp in, customer identified, catalogue searched, basket priced against real margin rules, owner approves on the dashboard, the graph resumes from its checkpoint, and the associate and customer both hear about it. The observability stack — OpenTelemetry into Prometheus, Grafana and Jaeger — is running, and the honest gaps are documented rather than hidden.

The backlog is mostly about the parts I deliberately left simple: true token-level streaming of the pause, multi-item bargaining across turns, and margin optimisation — suggesting the discount that clears the threshold rather than only refusing the one that does not. And the supervisor made routing probabilistic, which is a real cost. It is mitigated by strict structured output and a temperature-0 policy, but “two identical messages may route differently across runs” is a sentence that deserves to be measured rather than asserted, and that measurement is the next piece of work.

If you are building something in this space — or you want to argue with any of the decisions above — I would genuinely like to hear it. You can find the rest of my work on GitHub or reach me at hello@kavindunirmal.com.

Until next time, keep building.

Co-developers

Interested in working together?

I’m always open to new opportunities and collaborations. Whether you have a project in mind, want to chat about potential roles, or just want to connect, feel free to reach out!

Get in Touch

Let’s Connect!

I’m active on LinkedIn and GitHub, where I share my projects and insights. Feel free to follow me to stay updated on my work and connect with me!


Kavindu Nirmal Delpachithra | 2026 | Sri Lanka