Scoring Trust Without a Survey: Rating Credibility from Public Signal Alone
You can rate a crypto project's trust-risk from public signals alone, and make the number defensible instead of a vibe, because three research traditions supply the backbone.
A founder asks the reasonable question. You looked at our project for a day and gave it a 62 out of 100. Where did 62 come from? If the honest answer is “it felt like a 62,” the number is theater, and a skeptical reader is right to bin it. Any trust score that can’t survive that question is a vibe wearing a lab coat.
The audit this desk runs produces exactly such a number, so the question isn’t rhetorical. It has an answer, and the answer is that you can rate a project’s trust-risk from public signals alone, without ever running a survey, and make the rating defensible rather than felt. Three mature research traditions supply the backbone. Web-credibility science tells you what to look at and how the untrained eye mis-judges it. Signaling theory tells you which signals to believe. Psychometrics tells you how to prove the whole thing isn’t arbitrary. Put them together and a checklist becomes an instrument.
What the score is actually estimating
Start with the thing being measured, because a number that can’t name its own target is already lost. The score estimates one latent quantity: the probability that a project’s operators are both competent and benevolent toward the people holding the token. Competence and benevolence are not the desk’s invention. They are the two dimensions the Hovland–Yale group established as the core of source credibility three-quarters of a century ago (Hovland, Janis & Kelley, 1953), later joined by a third, goodwill, meaning perceived good intent toward the receiver (McCroskey & Teven, 1999). Trustworthiness, expertise, goodwill. A trust score is a bet on those three, read off a project instead of a person.
Web-credibility research already solved the harder half of this problem, rigorously and at scale, and most people in crypto have never heard of it. B.J. Fogg’s Stanford work synthesized studies of over 4,500 participants into a working theory of how people judge credibility from surface signals (Fogg, Soohoo, Danielson, Marable, Stanford & Tauber, 2003). His Prominence–Interpretation Theory (Fogg, 2003) says a credibility judgment needs two separate things to happen: an element has to be noticed (prominence), then it has to be judged (interpretation). That two-step is quietly the design brief for any rubric. A good rubric does the two jobs the naked eye does badly. It forces prominence, making sure a rater actually registers the signals that matter rather than the flashy ones, and it standardizes interpretation, fixing how each signal is scored so the result doesn’t drift with taste.
The same body of work also tells you which signals are weakest, and it’s uncomfortable. Fogg’s large study found that “design look” was the single most-mentioned factor when ordinary people rated a site’s credibility, cited in almost half of all comments (Fogg et al., 2003). People judge a site by how it looks. That is precisely the cheap signal a scam optimizes first: a polished landing page, a slick whitepaper, a crowded Telegram. Fogg’s four-part typology of credibility gives the vocabulary for demoting exactly that. Presumed credibility rests on general assumptions, surface credibility on first-impression inspection, reputed on third-party endorsement, and earned on repeated interaction over real time (Fogg, 2003, Persuasive Technology). The method’s job is to demote presumed and surface, and to elevate earned and the verifiable kind of reputed. Sundar’s MAIN model extends the point to machine cues: technological affordances throw off their own credibility heuristics (Sundar, 2008), which is the theoretical license for treating an on-chain, machine-verifiable fact differently from a human claim. The chain doesn’t have a marketing department.
Which signals to believe
This is the load-bearing pillar, and it comes from economics and evolutionary biology rather than design. A signal carries information only to the degree that faking it is costly for a bad actor. Michael Spence won a Nobel for the version economists use: education works as a labor-market signal not because school teaches the relevant skill but because it is differentially costly, cheaper to acquire for the high-quality candidate than the low-quality one (Spence, 1973). Amotz Zahavi reached the same place in biology with the handicap principle, the argument that a signal stays honest only when faking it is prohibitively expensive (Zahavi, 1975). Two fields, one conclusion, which is roughly the strongest evidentiary position an idea can occupy. Connelly and colleagues, reviewing three decades of signaling research across management, treat the honest-signal-requires-cost result as the settled core of the field (Connelly, Certo, Ireland & Reutzel, 2011).
That single principle sorts every public signal a crypto project emits into two piles. The sorting is the method’s spine, so it gets one format and holds it:
- Costly / hard to fake. A reputable third-party audit with its findings actually remediated. Source code verified on the block explorer, meaning the public can read the exact code that’s running. Ownership renounced (the deployer gives up the private key that controls the contract, so nobody can change it) or governed by a genuine multisig with a high signing threshold (several independent keys required to move anything, not one person’s wallet). Liquidity locked for a long horizon or burned outright (the trading pool can’t be yanked). A treasury that is transparent and time-locked. A doxxed team with a verifiable prior track record, meaning real people staking a reputation that can be destroyed. A multi-year operating history without incident, which is Fogg’s earned credibility and the one signal that can’t be manufactured because it needs the passage of real time. Each of these costs something real: forfeited control, surrendered exit options, destroyable reputation.
- Cheap / easy to fake. Follower counts. Testimonials. “Team size” claims. Advisor logos. Ambassador programs. A long glossy whitepaper. Paid listing boosts. Unaudited claims of an audit. None of these imposes a cost on a bad actor, so none of them survives contact with an adversary.
The crypto-specific evidence confirms the cheap pile is exploited exactly as theory predicts. Ante, Sandner and Fiedler, studying blockchain token offerings, document how readily social-media and human-capital signals were deployed to raise money in a market with weak verification (Ante, Sandner & Fiedler, 2018). The asymmetry is the part worth internalizing: costly signals are informative when present, while cheap signals are informative mainly when contradicted by the chain. “Community-owned,” set against a holder map showing seventy percent of supply in one connected cluster, is worth more than any number of ambassadors. So a trust-risk score should almost never let a cheap signal raise the number, and should let its contradiction lower it.
One honest complication, because the desk cites against itself when the record demands it. Fisch, analyzing 423 token offerings, found that technical whitepapers and high-quality source code did predict how much a project raised, while patents did not (Fisch, 2019). Substance in the documentation is informative; gloss is not. That is not a hole in the signaling account, it’s a refinement of it. The costly part of a whitepaper is the verifiable, falsifiable commitment buried inside it, not the page count.
Two columns, one axis. Every scored indicator placed by cost-to-fake, from a burned liquidity pool and a renounced contract at the expensive end to follower counts and advisor logos at the free end. The height of each is its weight in the score. The picture makes the argument the prose does: the number leans on the expensive side of the line.
Making it not arbitrary
A rubric with the right signals is still just an informed opinion until it clears two bars psychometrics has demanded of any measurement instrument for decades. It has to be reliable, meaning two trained people scoring the same project independently land in the same place. And it has to be valid, meaning the number actually measures trust-risk and actually predicts bad outcomes.
Reliability is where “isn’t this just subjective?” gets answered concretely, and the answer is borrowed wholesale from content analysis, the discipline that has spent fifty years teaching humans to code messy material the same way twice. You write a codebook: an operational definition and a decision rule for every indicator, with a worked example. You train at least two coders on a pilot sample kept separate from the real set, and iterate the rules until agreement stabilizes (Neuendorf, 2017). You double-code a random subsample. Then you report a chance-corrected agreement statistic, Krippendorff’s alpha, per indicator rather than as one flattering average, because an eighty-percent average across five indicators can hide one indicator sitting at zero. Krippendorff’s own guidance is strict, and appropriately so for anything published: rely on data at alpha of 0.800 or above, treat the 0.667-to-0.800 band as fit only for tentative conclusions, and discard the rest (Krippendorff, 2004). If two trained strangers reach 0.80 on an indicator, that indicator is inter-subjectively reproducible, which is the working definition of a measurement rather than a taste.
Validity is the harder and more interesting bar, and it has a decisive test. Content validity is the bookkeeping: every indicator maps to a theoretical pillar, a credibility dimension or a signaling cost, and you tabulate the mapping so a critic can check it (Cronbach & Meehl, 1955). Construct validity asks the score to correlate with independent trust ratings without collapsing into any one of them. But the test that settles the argument is predictive validity: assemble a labeled set of projects with known outcomes, the rugs and exploits and quiet abandonments alongside the survivors, score them blind to outcome, and measure whether the score separates the two. The standard instrument is the ROC curve and its area (Fawcett, 2006), where 0.5 is a coin flip and higher is real discrimination. Because rugs are a rare event relative to the flood of tokens, you also report a precision-recall curve, which stays honest under that imbalance where ROC alone flatters. The machine-learning literature shows features of exactly this kind already discriminate near-perfectly on historical data. Xia and colleagues characterized scam tokens on Uniswap at scale and built detectors from on-chain features (Xia et al., 2021); Mazorra, Adan and Daza reported a model separating fraudulent from legitimate tokens with accuracy above 0.99 on a large labeled set (Mazorra, Adan & Daza, 2022). That sets a realistic ceiling and a warning in the same breath: in-sample near-perfection always overstates live performance, so calibration on held-out data is the number that matters, not the headline on the training set.
Why you don’t need clever weights
The last objection is about the weights. Seven dimensions, each scored and combined, and the obvious worry is that the weighting is the point where taste sneaks back in. Which dimension is worth twenty-two percent and which is worth eight, and who decided?
The reassuring finding is decades old and solid enough to lean on. Robyn Dawes showed that equal or unit weights on well-chosen predictors often forecast as well as, sometimes better than, statistically optimized weights (Dawes, 1979). Optimized weights overfit the sample they were tuned on; equal weights don’t, and they carry the bonus of being transparent and hard to game. This is the single most useful license the method has. Sensible, near-equal weights on the right indicators are defensible by default, tilted only toward the dimensions carrying the most hard-to-fake signal. Code-and-security and on-chain integrity, the two dimensions dominated by machine-verifiable costly signals, together earn the largest share, a deliberate lean away from documentation gloss and community metrics. Empirically fitted weights become the upgrade path, adopted only once enough labeled outcomes exist and only if they beat equal weights on data they weren’t fitted to. Dawes predicts that gain will be small. Ship the equal weights, and treat the fancy version as something to earn.
One more mechanism keeps the arithmetic honest against a real failure mode. A weighted average can let a pile of green signals drown one fatal red one, and in this domain a single fatal signal should end the conversation. So the method applies critical-flag overrides: a honeypot pattern in the contract, an unlocked liquidity pool sitting in a wallet that also holds mint power, more than half the supply in one connected non-exchange cluster. Any one of these caps the total in the bottom tier regardless of how clean everything else looks. It encodes the practitioner’s rule that one red flag outweighs a dozen green ones, because the costly-signal logic says the red flag is the one that’s expensive to fake and therefore the one telling the truth.
What the number can and can’t say
None of this makes the score a safety guarantee, and pretending otherwise would betray the whole method. Audited projects have still rugged. Scanners get gamed by attackers who code around the known checks. A contract that looks renounced can hide a swappable proxy. The number is a lower-risk verdict, not a promise, and it decays: an upgradeable contract or a tax switch or a shift in holder concentration can move it after the scan, so it gets re-run on a schedule and on events. The base rates are genuinely brutal, with some token populations turning out fraudulent at rates near ninety-eight percent, which means even a strong model produces false readings and has to communicate its uncertainty rather than imply a precision it doesn’t have. And the instrument measures trust-risk, not investment merit. A trustworthy project can still fail on its fundamentals. A low score flags governance and security risk, never a price.
What the method does buy is the thing the founder asked for at the start. When the number comes back 62, there is a codebook that says why, a second trained rater who reached roughly the same 62, a mapping from every indicator to a published theory of credibility or a measurable cost of faking, and a labeled history against which the whole scale was calibrated. The 62 is reproducible, traceable, and falsifiable. That is the entire difference between an instrument and an opinion, and it’s the reason the score can go in a report a skeptic reads with money on the line.