NYC MyCity: a government chatbot that advised breaking the law

The Markup found in 2024 that New York City's AI chatbot for business owners gave answers contrary to city law, including on tenants with housing vouchers.

Year
2024
Jurisdiction
New York City, United States
Sector
Public sector: business guidance
Evidence base
Secondary sources
Incident record
OECD AI Incidents and Hazards Monitor · AIID 714
Harm
Reputational damage from harmful or unlawful public outputs · Legal liability for what the system tells customers

What happened

In October 2023 New York City announced an AI-powered chatbot to help business owners navigate government; it runs on Microsoft's Azure AI services 1.

The Markup reported on 29 Mar 2024 that in its testing the bot said landlords did not have to accept tenants with housing vouchers, although source-of-income discrimination is illegal in the city, and told a user they could take a cut of workers' tips. On the voucher question the bot once told a reporter that landlords did have to accept vouchers, then told ten separate staffers that they did not 1.

A city spokesperson said the chatbot was a pilot that would improve, had already given thousands of people accurate answers, and disclosed its risks to users 1. The OECD AI Incidents Monitor and the AI Incident Database both record the case 23.

Failure mode

A generative system answered legal questions in a domain where a wrong answer invites unlawful conduct, with no check that answers matched the law, and with answers that changed from one asking to the next. A disclaimer told users about the risk; nothing reduced it.

Which control would have caught it

An eval gate built on a legal-question set written with the agencies that enforce each rule (tips, housing, cash acceptance) measures the error rate before launch and sets a bar to clear. A runtime guardrail that grounds answers in official sources and refuses when it cannot cite one, and a red-team pass over the questions the public will ask, close the gap between a pilot label and a public service.

Patterns: Eval Gate in CI · Runtime Guardrail · Adversarial Red-Team Suite

The evidence that would have existed

What an auditor could have read, and the stack layer that produces it.

  • L3 Eval report on the legal-question set with the launch threshold and error rate per topic
  • L3 Consistency results: the same question asked many times, with the spread of answers
  • L4 Guardrail logs with the official source cited for each answer, or the refusal
  • L3 Red-team findings, each closed or formally accepted before launch

Obligations it touches today

As of 2026-09-24. Mappings are illustrative, not a claim of conformity.

  • NIST AI 600-1 Confabulation The generative-AI profile lists confabulation, confidently stated but false content, among its twelve risks 4.
  • EU AI Act Art. 50(1) For an EU deployment, people must be told they are dealing with an AI system 5. The duty is about disclosure, not accuracy, which is why the eval gate matters more than the banner.

System boundary

The MyCity chatbot as the city launched it: a generative model on Microsoft's Azure AI services, the questions business owners typed and the answers it returned about city rules on housing, employment and running a business 1. The agencies' rules sit outside the system as its reference; the people who act on an answer sit outside it too, and that is where the harm lands.

Control assumptions

What the controls below take for granted. Challenge any of them.

  • An answer about the law is correct only if it matches the rule the enforcing agency applies; fluency and confidence are not evidence of either.
  • The same question can get different answers: in The Markup's testing the voucher answer changed from one asking to the next 1, so one passing run of a test set shows little.
  • A disclaimer describes the risk without reducing it; the error rate falls only through a check before launch and a guardrail at the point of output.

Controls by moment

Preventive

Before the failure: design choices and release gates.

Detective

While it happens: what notices it.

Responsive

After it: what contains it and feeds the fix back.

Evidence requirements

The evidence each control must leave, written as acceptance criteria.

  • An eval report on a versioned legal-question set, written with the agencies that enforce each rule, shows the error rate per topic below the launch threshold for the version that goes live.
  • Consistency results show each question in the set asked many times, with the spread of answers recorded and within the threshold: a validity check that the eval measured the answers people get, not one lucky run.
  • Every answer in the guardrail log cites the official source it rests on, or is a refusal.
  • Every red-team finding is closed, or formally accepted with an owner, before launch.

Draft control specifications from the open control profiles, open for technical review.

Open questions

  • What error rate on a legal-question set is low enough to launch a public service that answers questions about the law, and who signs that threshold off?
  • How should an eval gate score a question that is answered correctly on one asking and wrongly on the next?

Sources

  1. [1] NYC's AI Chatbot Tells Businesses to Break the Law (investigative testing of the MyCity chatbot). The Markup. 2024-03-29. https://themarkup.org/news/2024/03/29/nycs-ai-chatbot-tells-businesses-to-break-the-law (verified: secondary)
  2. [2] OECD AI Incidents Monitor: NYC MyCity Chatbot Gives Dangerous, Illegal Advice to Businesses. OECD.AI. 2024-03-29. https://oecd.ai/en/incidents/2024-03-29-3dce (verified: primary)
  3. [3] AI Incident Database, Incident 714: Microsoft-Powered New York City Chatbot Advises Illegal Practices. Responsible AI Collaborative. 2026. https://incidentdatabase.ai/cite/714/ (verified: primary)
  4. [4] Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1; confabulation listed among twelve generative-AI risks). NIST. 2024-07-26. https://doi.org/10.6028/NIST.AI.600-1 (verified: primary)
  5. [5] EU AI Act Art. 50 (transparency; systems that interact directly with natural persons must be designed so that those persons are informed they are interacting with an AI system). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_50 (verified: primary)