When Social Conflict Becomes a Recommendation Signal: Why Recommendation Algorithms Cannot Stay Neutral

Table of Contents
August 17, 2026, evening, a gathering of people occurred downstairs at the Yangtze International office building in Chongqing’s Nan’an District. Police reported that a few individuals at the scene had a verbal altercation, which escalated into splashing water and chasing. Police subsequently stopped the conflict, dispersed the crowd, and provided legal education to several individuals involved [1].
This location is the office and training facility for the idol management company Times Fengjun in Chongqing, and young fans have long waited or checked in downstairs. Media analysis of on-site videos and public information indicates that before and after the incident, some male streamers also arrived to broadcast live, shouting internet memes that fans found annoying, leading to arguments and physical altercations between the two sides [2]. These details mainly come from media reports and online videos and should not all be considered facts confirmed by the police.
For the participants on site, it was an event of provocation, counterattack, and loss of control; for the live stream, it was simultaneously a constant stream of viewing, commenting, and sharing. The more intense the fan reaction, the more dramatic the live stream; the more onlookers, the more motivated the streamer to continue filming. This does not mean that recommendation algorithms created the on-site conflict, nor is there public data proving exactly which live streams were amplified by platforms. However, traffic incentives have already begun to influence participants’ behavior patterns. This street dispute, initially centered around fan check-ins and male streamers’ sensationalist curiosity, quickly became labeled with gender tags like “polarized fans” and “misogynistic provocation” after entering online platforms, escalating into a public event framed as male-female antagonism.
The conflict downstairs at Yangtze International cannot represent Chinese men or women as a whole, but it vividly demonstrates how social contradictions are transformed into content. Real-world gender dissatisfaction provides the emotional basis for conflict; streamers convert these emotions into performances for an audience; and user reactions become signals that platforms can record. While platforms do not create fundamental problems like housing, bride price, childbirth, and household division of labor, they can change the visibility of different experiences and feed these altered behaviors back into the next round of recommendations.
Therefore, to understand modern male-female antagonism, one must both explain why costs and responsibilities are not evenly distributed in reality and inquire into how platforms participate in shaping people’s understanding of these contradictions, as well as what responsibilities engineers should bear.
I. The Dual Logic of Traffic: How Offline Conflicts Mutate into Online Business #

Facts confirmed by the police include a gathering of people, verbal altercations, splashing water, and chasing. As for who initiated the provocation, who specifically came to the scene for traffic, and the details of injuries and handling circulating online, accounts from different sources are not entirely consistent. Analyzing platform mechanisms does not require filling in unconfirmed facts for any party, nor should on-site participants be regarded as representatives of an entire gender.
What is truly noteworthy about this case is how live streaming changed the conflict’s benefit structure. Arguments originally confined to the scene, transmitted online via phone cameras, could bring in new viewers with every provocation and counterattack. Live stream comments, dwell time, and shares would also instantly inform streamers which actions best maintained attention. Streamers do not need to agree with a particular gender stance; they might continue to play the provocateur simply because the audience reacts.
Fans’ counterattacks also enter this cycle. For those involved, cursing back or splashing water might be seen as defending their idol and their own dignity; for content dissemination, more intense reactions are more likely to become the climax of the next video. Counterattacks might force provocateurs to stop, or they might provide them with new content. Platforms can record interaction intensity but struggle to understand the underlying disgust, fear, and compelled responses.
This is not proof that “men are inherently hostile towards women” or that “female fan engagement necessarily leads to irrationality.” The conflict, initially about idols, live streams, and public order, was only gradually generalized into male-female antagonism after entering the internet. Specific individuals and behaviors were compressed into group labels, and complex scenes were edited into short videos suitable for taking sides.
Technology is not a passive bystander outside society here. The reason why netizens under the videos split into two major camps engaging in fierce verbal attacks, seemingly evaluating the specific altercation downstairs at the office building, is in fact to use it as an excuse to project their accumulated grievances regarding gender interests and long-term anxieties from real life onto this specific incident. Platform algorithms capture these projected emotions through recommendation systems to gain clicks and dwell time. However, to explain where this “real-world grievance” that can continuously fuel the conflict comes from, we must look beyond the internet’s live streams and enter the most core and widespread male-female cooperative contract in reality—marriage.
II. The Cost Mismatch in Marriage: Incommensurable Bills Borne by Men and Women #

In online conflicts concerning intimate relationships, the grievances and hostilities most frequently vented often revolve around bride price, wedding housing, childbirth, and the division of household labor. These polarized debates are not unfounded; they fundamentally reflect the sharply different costs and risks borne by men and women when entering and maintaining the contract of marriage.
When men discuss marriage costs, they most often cite property, bride price, wedding, car, and stable income. These expenses largely occur before marriage or in its early stages, are clearly quantifiable, and are often publicly compared, thus resembling an ever-increasing admission ticket. A study using Chinese urban data from 2000 to 2005 found that, within its model and sample, for every 1% increase in housing prices, the first marriage rate decreased by approximately 0.31% [3]. Older urban data cannot represent all regions today, nor does it imply that housing pressure is borne solely by men, but it supports a limited conclusion: once housing becomes a prerequisite for marriage, rising housing prices can elevate the barrier to marriage.
Women’s concerns about costs are more distributed throughout the duration of the marriage: childbirth brings physical risks, childcare affects career continuity, caregiving consumes time, and prolonged absence from the workplace impacts income, promotions, and retirement benefits. Research on the Chinese labor market has also observed a “Motherhood Penalty”: becoming a mother has a negative impact on women’s labor outcomes; grandparental care may alleviate some pressure, but the severity of the motherhood penalty has increased over the past few decades [4]. While research cannot dictate everyone’s career trajectory, it indicates that reproductive risks can indeed translate into economic costs.
The National Bureau of Statistics’ 2018 National Time Use Survey showed that residents’ average daily housework time was 1 hour and 26 minutes, with men spending 45 minutes and women spending 2 hours and 6 minutes [5]. While national averages cannot represent every household, they indicate that unpaid labor is generally still not equally distributed.
These two types of costs are difficult to directly add up. Property, bride price, and wedding expenses can be immediately priced, but childbirth, caregiving, and career interruptions involve future risks. The market turns housing, education, and salary into comparable mate selection criteria but does not provide a recognized price for housework and emotional support. Bride price also has different meanings. Wan’s study of over 100,000 judicial documents distinguished two common frameworks: compensation between families for women’s labor and childbearing, and intergenerational property gifted by parents to the newly married couple [6]. If both parties understand the same sum of money as different arrangements from the outset, disputes over repayment responsibility naturally arise after a relationship breaks down.
The problem is not which party “overall suffers more,” but that marriage allocates different costs to different people without a mutually agreed-upon method of measurement. Men see the explicit costs already paid, while women worry about long-term risks yet to materialize, and the argument thus becomes two unacknowledged bills.
III. Structural Lag in the Social Contract: When Rights Awareness and Family Division of Labor Disconnect #

Compared to the previous generation, today’s Chinese women have more complete educational and professional opportunities. The one-child policy had complex effects on this. Anthropologist Vanessa Fong’s research on urban only-daughters found that when daughters no longer competed with brothers for family resources, some families concentrated educational investment and upward aspirations on their daughters [7]. This study cannot represent all regions and classes, but it reveals how population policies unexpectedly altered internal family resource allocation.
Meanwhile, a relatively high birth sex ratio left another structural impact. The Seventh National Population Census showed that the national total population sex ratio was 105.07, and the birth sex ratio in 2020 was 111.3, which was a decrease of 6.8 from 2010 but still relatively high [8]. Population ratios do not determine whether individuals can marry, but in certain age groups, regions, and communities, demographic imbalance can indeed intensify marriage competition for men.
The problem is that the adjustment in rights awareness, professional opportunities, and family responsibilities for both sexes is not synchronized. After women entered universities and offices, society did not automatically reallocate reproductive and caregiving responsibilities; after men lost some traditional family authority, they did not simultaneously shed the expectation of providing housing, income, and security. More and more people emphasize individual rights, yet may still demand that their partners fulfill traditional obligations.
When income growth and employment were stable, families could absorb contradictions with future gains; when income expectations weaken and employment uncertainty rises, both men’s expectation of asset appreciation and women’s trust in long-term security decline. Economic pressure not only makes people “unable to afford marriage” but also weakens the basis for fulfilling old marriage contracts. The structural source of modern male-female antagonism lies in the fact that individual rights, professional opportunities, population structure, and family responsibilities have changed independently without collectively forming a new marriage contract. Once on the internet, these contradictions then transform into quantifiable behavioral signals.
IV. Anger Doesn’t Equal Liking, But Algorithms May Not Differentiate #

Recommendation systems struggle to directly measure whether users understand other groups better, obtain more accurate information, or are more capable of establishing stable relationships. Platforms more easily observe clicks, dwell time, complete views, comments, shares, and return visits. These actions, which do not require active rating, are called implicit feedback, and engineering teams use them to infer interest and decide what to display next.
But behavior does not equal preference. Someone watching a “sky-high bride price” video for a long time might do so out of agreement, or out of anger, skepticism, or to prepare a rebuttal. Recommendation systems cannot directly judge whether content helps users understand society; they can only use click-through rates, watch time, and interaction rates as proxy metrics. Proxy metrics are not useless; the danger is when teams forget they are just approximations: the system might interpret anger as interest, debate as satisfaction, and inability to stop watching as a desire to see more.
This creates a feedback loop: users can only react to content they have already seen, and the system then feeds those reactions into the next round of recommendations. Conflict-laden content gains more exposure, making it easier to accumulate dwell time and comments; new interactions then prompt the model to further increase exposure. In my previous article, I discussed how user choices influence recommendations, and recommendations, in turn, determine what choices users have next. The “preferences” observed by platforms are actually the result of the combined action of existing inclinations, content supply, and exposure decisions.
A preregistered algorithm audit published in PNAS Nexus in 2025 compared Twitter’s engagement-ranked feed with its chronological feed. In political content experiments, engagement ranking amplified anger, partisan leanings, and out-group hostility more, while users did not necessarily like these tweets more afterward [9]. The study cannot directly represent gender topics on Chinese platforms but reveals a transferable mechanism: content most likely to provoke a reaction is not necessarily what users truly want to see after reflection.
Therefore, the statement that “algorithms merely cater to users” is incomplete. Users provide behavioral signals, but platforms decide what to record, how to weigh it, which content enters the candidate set, and how to handle negative interactions. These are all product and engineering choices, not natural laws.
V. The Cost of Feedback Loops: When Exposure Decisions Contaminate Bias Data #

People often understand algorithmic bias as a data quality issue: societal data is biased, and the model learns from it. This is only half true. Recommendation data comes from exposures already arranged by the platform: if content doesn’t enter the candidate set, users can’t click on it; if a certain viewpoint is ranked higher, it’s more likely to accumulate interactions; creators, seeing what gains traffic, will also adjust their subsequent production. The model learns not raw society, but a society shaped by cultural notions, real demands, historical exposure, and creator adaptation.
In male-female topics, this cycle is particularly evident. Housing pressure, household division, and reproductive risks generate dissatisfaction; content producers compress this dissatisfaction into labels like “gold digger” or “delusional men”; recommendation systems expand their visibility based on interactions; and repeated exposure makes it easier for users to perceive extreme cases as universal realities.
Users and creators also learn what kinds of expressions are more likely to elicit feedback. Brady et al., after analyzing Twitter data and conducting controlled experiments, found that the more positive feedback moral outrage received, the more likely participants were to express themselves that way in the future [10]. The study cannot be directly extrapolated to gender discussions in China but indicates that likes, shares, and comments do not just measure expression but also train expressers in return.
Content production is also influenced by competition for clicks. Robertson et al., using Upworthy’s large-scale randomized headline experiments, found that under the experimental conditions of that website, negative words in headlines increased click-through rates [11]. This doesn’t prove that all platforms reward negative content, but it explains why editing and labeling an “office building downstairs altercation” (which originally mixed fan culture with traffic games) with the polarized tag “male-female street war” is more suitable as a headline than reporting the procedural fact that “a verbal altercation led to splashing water and chasing on site”: procedural facts require readers to retain uncertainty, while conflict narratives immediately provide emotion and roles.
Algorithms thus alter the shape of conflict. Dissatisfaction regarding institutions, relationships, or specific experiences, after being tagged and repeatedly recommended, is more easily interpreted as an inherent nature of the other gender. Private dilemmas thus become group attributions.
The solution is also not to replace interaction ranking with users’ explicitly stated preferences. The aforementioned research found that doing so could reduce anger and out-group hostility, but it might also reduce users’ opportunities to encounter different viewpoints [9]. Engineering practice faces not an either/or situation, but how to make explicable trade-offs between relevance, satisfaction, diversity, public impact, and user autonomy.
Merely reducing similar content may not change group attitudes in the short term. A 2023 Facebook field experiment reduced the proportion of 23,377 participants’ exposure to like-minded political sources by about one-third. The adjustment reduced some uncivil content and content frequently spreading misinformation but did not significantly change the eight political attitudes measured in the study [12]. American political experiments cannot predict Chinese gender topics but remind us: a more diverse recommendation list does not mean users truly understand different experiences; intervention might also cause resistance, avoidance, or a shift to other channels.
VI. Personalized Persuasion: How Generative AI Erodes Public Consensus #

Traditional recommendation systems primarily determine which article or video users see; a single piece of content usually remains a public object that can be jointly verified. Generative AI, however, can select materials, adjust wording, arrange evidence, and generate complete causal explanations based on historical conversations, explicit preferences, and current questions. Personalization is no longer just “what to show whom,” but delves into “how the same event should be understood.”
Regarding this conflict, a system could answer three types of questions:
- Why do female fan groups irrationally mob passersby for their idols?
- Why do male streamers exploit curiosity and verbally provoke women to gain traffic?
- Which facts have been confirmed, and which are still online rumors or deliberate gender attributions?
The first two questions might cite real materials but would select facts favorable to their respective narratives; the third question requires the system to maintain uncertainty and source boundaries. If the model also knows a user’s gender views, emotional experiences, and linguistic style, it can choose the argument most easily accepted by that user. Different users might receive not just different viewpoints, but also different factual emphases, causal sequences, and attribution of responsibility.
Research indicates that Large Language Models (LLMs) can generate personalized persuasion content at scale. Matz et al.’s research found that information generated based on user characteristics can influence participants’ attitudes and behavioral intentions [13]; another 900-person experiment showed that in a set social issue debate, GPT-4, when given participants’ basic demographic information, was more effective at persuasion [14]. These experiments do not prove that generative AI has already exacerbated male-female antagonism in China, but they do show that models can adjust persuasive content based on user characteristics. When combined with long-term user profiles, recommendation systems, and interaction metrics, platforms might simultaneously control content selection, factual interpretation, and expression generation.
The same capability can also serve better goals. Argyle et al.’s designed AI conversational assistant does not choose a stance for participants but instead provides real-time suggestions for expressions that make the other person feel more understood. The experiment improved conversation quality, reciprocity, and tone but did not systematically change policy stances [15]. Another study had a “Habermas Machine” iteratively generate group statements based on individual opinions and subsequent critiques; among 5,734 British participants, people generally preferred the AI-generated versions [16]. While structured political experiments cannot prove that models can mediate gender conflicts on open platforms, they do suggest that models can rephrase attacks into specific claims, find common facts, and organize public statements while preserving disagreements. The key lies in whether the system optimizes for persuasion, interaction, or understanding.
Generative answers are also harder to verify. Besides “based on which report,” users also need to know what personal information, prompts, and ranking conditions the system used. Providing links alone is not enough; every verifiable claim should be supported by proximate sources. The ALCE benchmark decomposes generation with citations into answer quality, citation completeness, and citation correctness; a 2023 experiment found that on the ELI5 dataset, even the best-performing tested system had about half of its answers not fully supported by citations [17]. For public controversies, the system should also distinguish officially confirmed facts, claimant assertions, media renditions, and model inferences. When materials are insufficient, it should maintain uncertainty instead of filling evidence gaps with fluent text.
VII. Re-engineering the Objective Function: How to Operationalize Public Interest #

Discussing the social impact of algorithms does not demand that engineers decide for the public which view on marriage and relationships is correct. Ranking models cannot solve housing costs, childbearing burdens, and household division, nor should they suppress legitimate debate. However, “technology is just a tool” likewise does not absolve responsibility. The Association for Computing Machinery (ACM)’s code of professional ethics lists promoting societal well-being, avoiding harm, and prioritizing the public interest as fundamental responsibilities [18]. These responsibilities need to be integrated into objective functions, evaluation metrics, release processes, and organizational governance.
Engagement Volume Cannot Dictate Everything #
If systems only optimize for click-through rates, watch time, and comment counts, conflict-laden content may gain a structural advantage. Teams can simultaneously introduce signals like satisfaction surveys, topic fatigue, explicit preferences, and long-term retention, but any metric is merely an approximation. Before launch, boundaries between short-term interaction, long-term satisfaction, and negative feedback should be defined; after launch, if interaction rises but is accompanied by a drop in satisfaction, an increase in reports, or topic fatigue, a special review should be triggered instead of simply declaring the experiment a success.
Likes, saves, reports, angry comments, and repeatedly returning to comment sections do not express the same intent. While models struggle to accurately identify the psychological reasons behind each behavior, teams can still distinguish positive, negative, and ambiguously meaningful interactions, reduce the weight of clearly negative behaviors on long-term user profiles, and differentiate between temporary attention and stable preferences.
The product should also allow users to view, modify, and delete interest tags, tell the system “I’m just temporarily looking into this,” or temporarily reduce a certain topic. TikTok already offers recommendation resets, topic frequency adjustments, and keyword filtering [19]; Instagram also allows users to reset recommendations, letting the system re-evaluate based on new interactions [20]. These features return some control over profiling to users, but their effectiveness still needs evaluation: can users find and understand the settings, how long do resets last, and will the system still interpret angry interactions as long-term interests?
Beyond Individual Content, Consider the Entire Feed #
Recommendation quality is not just about clicks; it’s also about whether content is repetitive, sources are singular, extreme cases disproportionately displace more representative materials, and whether users can disengage from a particular topic. This does not mean mechanically dividing traffic equally between men and women. Teams can audit whether hostile content is escalating, whether out-group attacks are treated as high-value interactions, and whether exposure is overly concentrated, all by topic and user group. When predefined risk boundaries are exceeded, interaction weights should be adjusted, topic diversity increased, or exploration scope widened, followed by a re-evaluation of user experience.
Adjusting ranking is not just theoretical. Piccardi et al. conducted a preregistered field experiment with 1,256 participants during the 2024 US presidential election, using large models to identify and rerank posts expressing anti-democratic attitudes and partisan animosity. Lowering the ranking of such content alleviated participants’ hostility towards political out-groups, while raising it had the opposite effect [21]. X platform’s US political experiments cannot directly represent Chinese gender topics but provide strong causal evidence: platforms do not need to delete viewpoints; merely adjusting exposure order can also influence group attitudes.
Another path is to use cross-group consensus as a ranking signal. X’s Community Notes requires a supplementary note to be affirmed by users with typically different evaluation patterns [22]. However, this mechanism also has blind spots. Bouchaud and Ramaciotti, analyzing approximately 1.9 million notes and 135 million ratings, found that highly polarized content was less likely to achieve cross-group consensus and more likely to remain hidden [23]. Thus, signals that connect different groups can participate in ranking but cannot replace professional verification, risk identification, and human governance.
For public controversies, generative systems need to distinguish between officially confirmed facts, media renditions, online rumors, and model inferences. Police confirmation of people gathering, verbal altercations, splashing water, and chasing downstairs at Yangtze International does not mean police have confirmed that one gender systematically mobbed the other. The system should also retain source and version records, with highly controversial content defaulting to displaying its factual status; if unable to distinguish between rendition and fact, it should be flagged for manual review or diffusion restricted, rather than continued elaboration.
The National Institute of Standards and Technology (NIST)’s AI Risk Management Framework integrates governance, risk identification, measurement, and management throughout the system lifecycle; the generative AI accompanying framework also emphasizes human review, tracking, recording, and management oversight [24][25]. The purpose is not merely compliance, but to enable the system to explain, roll back, and correct errors when they occur.
Algorithm Responsibility Is Not Just an Engineer’s Business #
Recommendation goals are collectively determined by multiple parties: management chooses business models and risk tolerance, product teams define growth targets, algorithm teams select metrics and implement models, and governance, compliance, and audit teams oversee execution. Expecting an engineer to resist an entire organizational incentive structure based on personal ethics is neither realistic nor would it address the true locus of decision-making power.
Engineers still have irreplaceable professional responsibilities. They are most aware of how data is collected, how metrics are calculated, where models fail for certain groups, and are best positioned to point out that “rising interaction” might merely be a byproduct of escalating conflict. Social responsibility is not about choosing the right viewpoint for society, but about understanding what the system is optimizing for, who it might harm, and whether errors can be audited, explained, and corrected.
Truly effective algorithm responsibility must become an organizational capability: product reviews should record public risks and accountable individuals, model launches should evaluate group and long-term impacts, anomaly monitoring should cover more than just latency and click-through rates, and incident post-mortems should question whether the system amplified avoidable harm. Who has the authority to pause experiments, request rollbacks, or accept residual risks should also be clearly defined before launch.
Conclusion #

The conflict downstairs at Yangtze International needs to be handled by the police according to specific actions and laws. This article is more concerned with another layer of change: when the on-site provocations and counterattacks entered live streaming, they were no longer just disputes between individuals but also content capable of attracting views, comments, and shares. Platforms recorded interactions but did not necessarily understand why people interacted.
The asynchronous changes in individual rights, family obligations, population structure, and economic conditions provided a real-world foundation for the conflict. Recommendation systems turned dwell time, anger, and debate into interaction signals, reallocating the visibility of different experiences; generative AI then organized fragmented materials into comprehensive explanations for different users. Technology learns from society and, in turn, changes society. Engineers cannot solve all structural contradictions, but they can reduce avoidable misunderstandings and hostilities through prudent ranking, verifiable generation, assisted dialogue, and user control.
This is not asking technology to maintain a fictitious absolute neutrality, but asking designers to acknowledge that data selection, metric definition, and ranking decisions all allocate attention. When this allocation is sufficient to change the information environment for a large number of users, public interest becomes part of the engineering problem. Technology cannot achieve societal reconciliation, but it can provide conditions for more accurate, restrained, and sustainable discussions.
References #
[1] 重庆市公安局南岸区分局. 关于南滨路一演艺公司楼下人员聚集事件的警情通报. 2026. https://new.qq.com/rain/a/20260820A0CC3C00.
[2] 联合早报. 下午察:时代峰峻门外的饭圈风暴. 2026. https://www.zaobao.com.sg/news/china/story20260820-9548174.
[3] WRENN D H, YI J, ZHANG B. House Prices and Marriage Entry in China. Regional Science and Urban Economics, 2019, 74: 118–130. DOI: 10.1016/j.regsciurbeco.2018.12.001.
[4] MENG L, ZHANG Y, ZOU B. The Motherhood Penalty in China: Magnitudes, Trends, and the Role of Grandparenting. Journal of Comparative Economics, 2023, 51(1): 105–132. DOI: 10.1016/j.jce.2022.10.005.
[5] 国家统计局. 2018 年全国时间利用调查公报. 2019. https://www.stats.gov.cn/sj/zxfb/202302/t20230203_1900224.html.
[6] WAN Y. Between Money and Intimacy: Brideprice, Marriage, and Women’s Position in Contemporary China. Demographic Research, 2024, 50(46): 1353–1386. DOI: 10.4054/DemRes.2024.50.46.
[7] FONG V L. China’s One-Child Policy and the Empowerment of Urban Daughters. American Anthropologist, 2002, 104(4): 1098–1109. DOI: 10.1525/aa.2002.104.4.1098.
[8] 国家统计局,国务院第七次全国人口普查领导小组办公室. 第七次全国人口普查主要数据结果新闻发布会答记者问. 2021. https://www.stats.gov.cn/zt_18555/zdtjgz/zgrkpc/dqcrkpc/ggl/202302/t20230215_1904005.html.
[9] MILLI S, CARROLL M, WANG Y, et al. Engagement, User Satisfaction, and the Amplification of Divisive Content on Social Media. PNAS Nexus, 2025, 4(3): pgaf062. DOI: 10.1093/pnasnexus/pgaf062.
[10] BRADY W J, MCLOUGHLIN K, DOAN T N, CROCKETT M J. How Social Learning Amplifies Moral Outrage Expression in Online Social Networks. Science Advances, 2021, 7(33): eabe5641. DOI: 10.1126/sciadv.abe5641.
[11] ROBERTSON C E, PRÖLLOCHS N, SCHWARZENEGGER K, et al. Negativity Drives Online News Consumption. Nature Human Behaviour, 2023, 7: 812–822. DOI: 10.1038/s41562-023-01538-4.
[12] NYHAN B, SETTLE J, THORSON E, et al. Like-Minded Sources on Facebook Are Prevalent but Not Polarizing. Nature, 2023, 620: 137–144. DOI: 10.1038/s41586-023-06297-w.
[13] MATZ S C, TEENY J D, VAID S S, et al. The Potential of Generative AI for Personalized Persuasion at Scale. Scientific Reports, 2024, 14: 4692. DOI: 10.1038/s41598-024-53755-0.
[14] SALVI F, RIBEIRO M H, GALLIOTTO C, et al. On the Conversational Persuasiveness of GPT-4. Nature Human Behaviour, 2025. DOI: 10.1038/s41562-025-02194-6.
[15] ARGYLE L P, BUSBY E, GUBLER J, et al. Leveraging AI for Democratic Discourse: Chat Interventions Can Improve Online Political Conversations at Scale. Proceedings of the National Academy of Sciences, 2023, 120(41): e2311627120. DOI: 10.1073/pnas.2311627120.
[16] TESSLER M H, BAKKER M A, JARRETT D, et al. AI Can Help Humans Find Common Ground in Democratic Deliberation. Science, 2024, 386(6719): eadq2852. DOI: 10.1126/science.adq2852.
[17] GAO T, YEN H, YU J, CHEN D. Enabling Large Language Models to Generate Text with Citations. Proceedings of EMNLP 2023, 2023: 6465–6488. DOI: 10.18653/v1/2023.emnlp-main.398.
[18] ASSOCIATION FOR COMPUTING MACHINERY. ACM Code of Ethics and Professional Conduct. 2018. https://www.acm.org/code-of-ethics.
[19] TIKTOK. More Ways to Discover New Content and Creators You Love. 2025. https://newsroom.tiktok.com/en-us/more-ways-to-discover-new-content-and-creators-you-love.
[20] META. Reshape Your Instagram With a Recommendations Reset. 2024. https://about.fb.com/news/2024/11/introducing-recommendations-reset-instagram/.
[21] PICCARDI T, SAVESKI M, JIA C, et al. Reranking Partisan Animosity in Algorithmic Social Media Feeds Alters Affective Polarization. Science, 2025. DOI: 10.1126/science.adu5584.
[22] X. Note Ranking Algorithm. Community Notes Guide. 2026. https://communitynotes.x.com/guide/en/under-the-hood/ranking-notes.
[23] BOUCHAUD P, RAMACIOTTI P. Community Notes Undermoderate Polarizing Content by Design, Creating Risks in Electoral Processes. Science Advances, 2026, 12(27): eaee6932. DOI: 10.1126/sciadv.aee6932.
[24] NATIONAL INSTITUTE OF STANDARDS AND TECHNOLOGY. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1, 2023. https://doi.org/10.6028/NIST.AI.100-1.
[25] NATIONAL INSTITUTE OF STANDARDS AND TECHNOLOGY. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1, 2024. https://doi.org/10.6028/NIST.AI.600-1.