A vendor demo is not a product tour. It is the only hour in the entire evaluation where you can make a vendor prove something in real time, and most buyers spend it watching slides.
The difference between a useful demo and a wasted one is the question format. “Do you support rule configuration?” gets a yes from every vendor in the category. “Have someone on your team build this rule now, and we will time it” gets you information nobody can prepare a slide for.
This guide covers how to run the session, the questions that actually separate platforms, and what a strong versus weak answer sounds like for each.
Before the demo: five ground rules
Send your agenda in advance and hold them to it. Vendors default to a fixed presentation. Tell them what you want demonstrated and in what order, and say plainly that you will skip the company overview.
Bring the people who will use it. Your BSA officer or lead analyst will notice in ten minutes what a procurement lead will not notice in three meetings. Let them drive part of the session.
Bring your own material. A sample data file, your risk assessment, your investigation SOP, and a rule you actually want to build. A demo on vendor data shows you what the vendor chose to show you.
Ask for numbers, not adjectives. “Fast”, “flexible”, and “AI-powered” are now universal in this category and differentiate nothing. Every question below is designed to produce a figure, a timed action, or a document.
Take notes on what was demonstrated versus what was described. These are different claims, and the gap between them is where most post-purchase disappointment lives. A useful convention: score anything merely described as a maybe, and anything shown live as a yes.
Functional coverage
Ask: “Which of these are native to your platform, and which come through a partner?” Run the list out loud: identity verification, customer risk scoring, transaction monitoring, watchlist screening, case management, regulatory filing.
Weak answer: a capability grid where everything is a green tick. Strong answer: clear separation of native from partner-delivered, with partners named and told you what is contracted separately.
Ask: “Show me what happens when a customer’s risk rating changes.” You are looking for whether monitoring adjusts automatically or whether someone has to remember to update a rule.
Ask: “Does a screening hit and a monitoring alert on the same customer land in the same case?” If not, your analyst assembles that picture manually, on every alert, forever.
Configuration
This is the criterion where demos are most misleading, because a solutions engineer using their own software daily makes anything look easy.
Ask: “Have a non-technical person on your team build this rule, right now.” Give them a real rule from your own program. Time it. Then ask what a compliance analyst at a customer institution would take.
Weak answer: the rule gets built by a solutions engineer with a developer console open, or the vendor offers to follow up with a recorded video. Strong answer: built live by someone non-technical, in minutes, in the same interface you would use.
Ask: “Do configuration changes carry professional services fees or change-request charges?” Get it in writing
This is the single most useful commercial question in the evaluation and it is almost never asked. Some contracts permit unlimited self-service configuration. Others route every threshold change through a paid change request with a lead time. The difference shows up in year two as a budget line nobody forecast, and as a control environment you do not actually control.
Ask: “Can we build rules using our own data fields, or only your predefined parameters?”
Monitoring and detection
Ask: “What is your p99 latency under production load?” Not average. Average hides the tail, and the tail is where your peak-hour transactions live.
Ask: “What throughput do you sustain out of the box?” Rarely asked and quietly important if you have promotional or payday peaks.
Ask: “What is your uptime, and what is the URL of your status page?” Then open it during the call.
Weak answer: an uptime percentage with no public page behind it. Strong answer: a figure and a live page you can check yourself, plus a straight answer on which number is contractual if their materials quote more than one.
Ask: “Can you stop a transaction before it settles, or only flag it afterwards?” On instant rails, detection after settlement is forensics.
Ask: “Do real-time, post-event, and batch monitoring use the same rule logic?” If each mode needs its own rules, you pay that cost permanently.
Alert quality
Ask: “Project our alert volume at our transaction count, and show me the working.”
This is the most valuable question in the entire evaluation, because alert volume converts directly into headcount, and headcount dwarfs the licence fee. Two vendors quoting identical fees can differ by several analysts in real cost
Weak answer: a false positive reduction percentage offered in place of a projection. Strong answer: a projection based on your transaction profile, with the assumptions visible and challengeable.
Ask: “Backtest a rule against sample historical data and show me the projected alert volume before it deploys.”
Ask, of any reduction percentage they quote: “Reduction from what baseline, over what period, at what kind of institution, and achieved how?” Threshold tuning, automated clearing, and suppression are three different things. A vendor comfortable answering is telling you the number is real.
Investigations and case management
Ask: “Walk me through one case end to end.” Alert, evidence gathering, escalation, approval, decision, filing, submission receipt. Count how many systems the analyst touches.
Weak answer: a dashboard tour that never opens an actual case. Strong answer: a real case worked start to finish in one workspace, with the audit trail visible at each step.
Ask: “Create a new case status and an SLA timer, live, without engineering.”
Ask: “What proportion of alerts close without an analyst opening them, and what evidence is recorded for each closure?” Automated clearing is valuable only if it is documented. Undocumented auto-closure is a finding waiting to happen.
Ask: “Here is our investigation SOP. What can your platform do with it?” This question separates platforms that impose a workflow from platforms that adopt yours.
Screening
Ask: “Which lists are native, and which need our own subscription?” If you already pay a data provider, ask whether you can bring it.
Ask: “How often is list data refreshed, per list?” Push for a contractual figure rather than “continuously”.
Ask: “Change a matching threshold live and show me the effect on historical hits.”
Weak answer: a single sensitivity slider, or matching logic described as proprietary. Strong answer: named, independently configurable algorithms, with the ability to test a change against past hits before applying it.
Ask: “Which list version was in force on this date last year?” A specific, awkward question that tells you whether they retain list-version history. Examiners ask this.
Explainability and audit trail
Ask: “Reconstruct a decision from two years ago, on screen, showing the rule logic in force at the time.”
Do not accept a description of audit capability. Ask them to do it. You will need this under adversarial conditions eventually, and you want to know now whether it takes five minutes or a support ticket.
Ask, if AI participates in triage: “Show me an AI-cleared alert and everything behind it.” Evidence chain, confidence reasoning, and the specific model version used. “The model flagged it” is not something you can give an examiner.
Ask: “Export an audit log in the format we would hand to a regulator.”
Reporting
Ask: “File a SAR from inside a case and show me the submission receipt.” Then ask whether filing is by API or whether someone exports and re-keys into a regulator portal.
Ask: “Show me the CTR workflow.” Frequently a gap, and a meaningful ongoing cost if you are cash-heavy.
Ask: “Show me analyst throughput and SLA adherence without exporting to a spreadsheet.”
Integration, security, and implementation
Ask: “Do you have an existing integration to our core, and which customer is running it today?” Then ask to call them.
Ask: “How many engineering hours do you need from us?” A number, in writing. Not “minimal”.
Ask: “What is your median go-live in days, named to a customer at our volume?”
Weak answer: a range, or a timeline with no customer attached. Strong answer: a figure, a named comparable customer, and a written scope of what your team supplies.
Ask: “Send us your SOC 2 Type II report and ISO certificate.” Then verify the certificate with the certifying body and read the report for exceptions rather than checking the logo.
Ask: “During implementation, who calibrates the rules?” If the answer is you, factor that in.
Ask: “After go-live, do we get a named person or a ticket queue, and what is the response time commitment?” Small compliance teams have no internal escalation path, which makes this weigh more than it appears to.
Commercial
Ask: “What is the pricing model, and what triggers an increase mid-contract?”
Ask: “What changes commercially and operationally when we double volume or add a jurisdiction?” In writing. Ask specifically whether a new market is a configuration change or an implementation project.
Ask: “Are ongoing monitoring, screening, and each data type priced separately or bundled?” Itemised. Ongoing screening billed separately from initial screening is a common surprise.
What a complete set of answers looks like
The point of demanding this specificity is that some vendors can meet it. As a reference for the level to expect, here is how one platform’s published documentation answers the questions above.
Flagright. Monitoring runs at a published 200ms p99 API latency and 1,200 requests per second out of the box, at 99.998% uptime with a public status page, across more than 1.4 billion transactions monthly. Real-time, post-processing, and batch share identical rule logic, and transactions can be blocked before settlement.
On configuration, validated rule creation time is 60 seconds, roughly three minutes as measured by customers, and Flagright states that rule changes require no SQL, no engineering tickets, and no professional services fees. Rules backtest against 90 days of history and run in shadow mode against live traffic before promotion, which is the alert-volume question answered structurally rather than with a percentage. Reported false positive reduction is up to 83% from threshold optimisation specifically and 93% across broader AI tooling, two different scopes, which is exactly the distinction to demand of any vendor quoting a single number.
On investigations, case management is native with one centralised queue across sources, configurable statuses and SLA timers, no-code maker-checker and escalation routing, and an ontology view for multi-hop relationships. The SOP question has a direct answer: you upload your investigation procedure and the AI agent builds an investigation flow from it, in a stated 20 minutes, with control ranging from silent evaluation to full automation. Reported 77% of alerts auto-cleared with high confidence, 94% analyst agreement, and alert-to-outcome time of 4 minutes against a 38 minute baseline. A QA module reviews every case against your SOP rather than a sample.
On screening, matching algorithms are named and independently configurable, Jaro-Winkler, Levenshtein, transliteration, phonetic matching, DOB delta tolerance, tokenisation, and stopword filtering, with per-list and per-mode thresholds, simulation against historical match activity, and full list-version audit history.
On explainability, every rule change is written to an immutable timestamped audit log with every version preserved and one-click rollback. AI decisions carry an annotated transaction timeline, typology citation, confidence score with contributing factors, and a specific model version, exportable in JSON or Excel.
On reporting, SAR and CTR filing runs by API direct to FinCEN and to 70+ GoAML countries from inside the case, with narratives pre-filled and the submission receipt stored in the audit log.
On implementation and support, published average go-live is two weeks with rules calibration included, on API-first architecture with more than 100 native integrations including published connectivity to Jack Henry Symitar and Fiserv DNA. ISO 27001:2022 and SOC 2 Type II certified. Reported 6 minute average support response, 24/7 coverage, and a dedicated CSM.
Where the answers are less complete, and what to press on: pricing is not published, so the commercial questions above must be answered in your own quote process. The platform is cloud-native only. Screening list refresh is described as continuous without a per-list interval, so ask for it contractually. Uptime appears as 99.998% on product pages and 99.99% on security documentation, so ask which is contractual. And if you intend to buy monitoring alone while keeping existing case management, confirm standalone operation and third-party alert ingestion, since the platform’s design assumes a unified stack.
Other vendors commonly demoed alongside it, worth verifying directly against these same questions: Unit21 and Hawk AI for combined fraud and AML with no-code configuration, Napier AI for a production sandbox approach to rule testing, Lucinity for investigation workflow depth, ComplyAdvantage as a screening and data layer, Nasdaq Verafin for US banks and credit unions, and the enterprise suites from NICE Actimize, Oracle, and SAS where scale requires them and a longer implementation is acceptable.
The eight to insist on seeing live
If the session runs short, these are the ones that cannot be answered on paper:
- A non-technical person building and deploying a rule, timed.
- A backtest projecting alert volume before deployment.
- Alert volume projected at your transaction count, with the working.
- One case walked end to end, alert to submission receipt.
- A two-year-old decision reconstructed with the logic in force then.
- A matching threshold changed against historical hits.
- Your own SOP loaded into the platform.
- The status page, opened during the call.
A vendor who can show all eight without rescheduling is telling you something about the product. A vendor who can show two is also telling you something.

