reading customer reviews online, checking rating and comments

(Credit: © Song_about_summer - stock.adobe.com)

Why That Restaurant Review Might Make You Easier to Scam

In A Nutshell

  • Researchers built an algorithm that inferred hidden social connections on Yelp using nothing but public review activity, no friend lists required.
  • At a 20% false-alarm rate, the model correctly recovered more than 60% of real Yelp social ties in two U.S. cities.
  • In the Pennsylvania sample, modeled returns for a phishing campaign jumped from 109% to 1,098% as attack scale grew from 500 to 10,000 impersonation attempts.
  • Adding statistical noise to the data made smaller phishing campaigns unprofitable, though larger ones could still turn a profit.

Posting an honest product review seems harmless enough. New research suggests it can help reveal someone’s social connections to cybercriminals.

A study published in the journal Information Systems Research found that ordinary behavioral data, like how long a review runs, can reconstruct hidden social networks with startling accuracy. Researchers built a model that inferred social ties from patterns in review activity, including how long users’ reviews were, then showed that an attacker armed with those connections could turn spear-phishing, a scam where a hacker impersonates a trusted contact, into a far more lucrative enterprise.

Yan Leng of the University of Texas at Austin led the team, alongside Yijun Chen of the University of Melbourne and collaborators from Oxford, the Chinese University of Hong Kong, and the University of Sydney. Their work didn’t stop at identifying the danger. They also built and tested two methods platforms could use to make this kind of snooping much harder, without destroying much of the data’s usefulness.

Online Reviews Helped the Algorithm Recover More Than 60% of Yelp Social Ties

An economic idea underpins the approach: people’s choices are shaped by their peers. Someone who sees friends write long, detailed reviews might feel compelled to write a longer one themselves, or might feel there’s less left to say. Either way, that peer influence leaves fingerprints in the data.

Building on that logic, the team designed an algorithm called Homophilous Network Learning, or HNL. Rather than looking at who follows whom, HNL works from patterns in users’ review activity, including which businesses they reviewed and how long those reviews were. It then works backward, asking a simple question repeatedly: given how everyone behaved, what hidden web of social connections would best explain these patterns? Researchers needed a real platform to test against, one with public reviews and a known friend network to check their guesses, and Yelp fit that bill.

Researchers used a public Yelp dataset covering reviews in two U.S. metro areas, New Orleans and Pittsburgh. The Louisiana sample included 2,065 reviewers and 6,094 businesses, while the Pennsylvania sample covered 2,234 reviewers and nearly 16,000 businesses. Yelp publishes an actual “friends” network for its users, giving the team a rare chance to check guesses against ground truth.

Results held up well. At a setting where 10% of nonexistent connections were mistakenly flagged as links, the model still recovered about half of the actual Yelp connections in both cities. Allowing more false alarms, up to 20%, pushed that figure above 60%. Older approaches, including simple correlation analysis and a few competing computer models, all lagged behind at recovering real ties.

Findings like these fed into an economic model of cybercrime. Drawing on loss figures from the FBI’s Internet Crime Complaint Center, researchers modeled how much money a phishing campaign might net an attacker who correctly guesses a victim’s social circle. As impersonation accuracy rises, so does what the paper calls “attack precision,” and payoff scales up fast. In the Pennsylvania sample, modeled return on attack rose from 109% for 500 impersonation attempts to 1,098% for 10,000. In the model, even modest attack precision could make large-scale phishing highly profitable.

online reivews infographic
Your review history might reveal who your real friends are, handing cybercriminals a phishing shortcut. (Image by StudyFinds)

Adding Data Noise Could Make Smaller Phishing Campaigns Unprofitable

Diagnosing the risk was only half the job. Researchers then designed two competing methods to protect users while preserving most of the data’s value.

Both approaches rely on differential privacy, a technique that injects a calibrated amount of random statistical noise into data before release. Done right, this noise throws off pattern-matching algorithms while leaving broader trends, like average review lengths, mostly intact.

One method, the direct mechanism, adds Gaussian noise (random variation shaped like a bell curve) straight onto review-length data before publication. A second approach, the indirect mechanism, works one step further removed: it perturbs the hidden network the algorithm has already inferred, then uses that scrambled version to regenerate a protected set of review-length data using Laplace noise, a distribution with a sharp peak and long tails.

Tested against the same Yelp data, both methods reduced the algorithm’s ability to guess real social ties while preserving most of the usefulness of the review-length data. Under one setting, about three-quarters of the review-length values were unchanged, and 95% of all entries differed by roughly 15 words or less. Even that modest scrambling flipped the economics of an attack: for smaller campaigns, estimated returns dropped into negative territory, a losing proposition. Larger campaigns still saw profits fall sharply compared to unprotected data, even where they didn’t vanish completely.

Between the two, the indirect mechanism generally confused attackers more effectively, though the direct mechanism was simpler to deploy.

Could Amazon and YouTube Face the Same Social Data Leak? Researchers Say It Is Possible

Authors frame the work as a wake-up call for an entire category of platforms. Groupon, Amazon, YouTube, and any service where users publicly rate or comment could, in theory, be leaking the same hidden social information. Review length was simply the easiest signal to test on Yelp’s public data; ratings, purchases, comments, or likes could potentially carry similar signals elsewhere, though this study did not test them.


Disclaimer: This article summarizes findings from a peer-reviewed study and is intended for general informational purposes. It does not constitute cybersecurity advice for any specific platform, account, or individual, and readers concerned about their own online privacy should consult a qualified security professional.


Paper Notes

Limitations

This study’s real-world test relied on a single platform, Yelp, and two U.S. metro regions, so how well the findings generalize to other review or rating platforms with different user behaviors remains untested. The researchers also treated Yelp’s published friendship graph as ground truth for scoring accuracy, but that graph itself may miss real offline relationships or include connections that carry little practical weight, meaning the true accuracy of the method could be somewhat over or understated. The analysis further assumes a fixed, static snapshot of behavior rather than accounting for how relationships and posting habits shift over time, which the authors note as an open direction for future work.

Funding and Disclosures

Lead author Yan Leng’s work on the paper was supported by a U.S. National Science Foundation grant (IIS-2153468). The authors report no other disclosed conflicts of interest.

Publication Details

Titled “When Behavioral Data Betray Users: A Diagnostic and Protective Framework Against Social Interaction Leakages,” the paper is authored by Yan Leng (University of Texas at Austin), Yijun Chen (University of Melbourne), Xiaowen Dong (University of Oxford), Junfeng Wu (Chinese University of Hong Kong, Shenzhen), and Guodong Shi (University of Sydney). It was published in Information Systems Research, DOI: 10.1287/isre.2024.1469. An earlier working paper version is also available via SSRN (DOI: 10.2139/ssrn.3875878).

About StudyFinds Analysis

Called "brilliant," "fantastic," and "spot on" by scientists and researchers, our acclaimed StudyFinds Analysis articles are created using an exclusive AI-based model with complete human oversight by the StudyFinds Editorial Team. For these articles, we use an unparalleled LLM process across multiple systems to analyze entire journal papers, extract data, and create accurate, accessible content. Our writing and editing team proofreads and polishes each and every article before publishing. With recent studies showing that artificial intelligence can interpret scientific research as well as (or even better) than field experts and specialists, StudyFinds was among the earliest to adopt and test this technology before approving its widespread use on our site. We stand by our practice and continuously update our processes to ensure the very highest level of accuracy. Read our AI Policy (link below) for more information.

Our Editorial Process

StudyFinds publishes digestible, agenda-free, transparent research summaries that are intended to inform the reader as well as stir civil, educated debate. We do not agree nor disagree with any of the studies we post, rather, we encourage our readers to debate the veracity of the findings themselves. All articles published on StudyFinds are vetted by our editors prior to publication and include links back to the source or corresponding journal article, if possible.

Our Editorial Team

Steve Fink

Editor-in-Chief

John Anderer

Associate Editor