---
title: "NYC MyCity: a government chatbot that advised breaking the law"
description: "The Markup found in 2024 that New York City's AI chatbot for business owners gave answers contrary to city law, including on tenants with housing vouchers."
canonical: https://aigovernanceengineer.com/cases/nyc-mycity-chatbot
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22956197
version: "0.5.0"
updated: 2026-09-26
---

# NYC MyCity: a government chatbot that advised breaking the law

> The Markup found in 2024 that New York City's AI chatbot for business owners gave answers contrary to city law, including on tenants with housing vouchers.

- Year: 2024
- Jurisdiction: New York City, United States
- Sector: Public sector: business guidance
- Evidence base: Secondary sources
- Incident record: [OECD AI Incidents and Hazards Monitor](https://oecd.ai/en/incidents/2024-03-29-3dce) · [AIID 714](https://incidentdatabase.ai/cite/714/)
- Harm: [Reputational damage from harmful or unlawful public outputs](https://aigovernanceengineer.com/resources/harms#harm-reputational-harm) · [Legal liability for what the system tells customers](https://aigovernanceengineer.com/resources/harms#harm-liability-for-outputs)

## In short

In October 2023 New York City announced MyCity, an AI chatbot to help business owners navigate government. The Markup reported on 29 Mar 2024 that in its testing the bot said landlords did not have to accept tenants with housing vouchers, although source-of-income discrimination is illegal in the city, and told a user they could take a cut of workers' tips; on the voucher question its answer changed from one asking to the next. The harms are reputational damage and legal liability for what the system tells people. The failure mode is a generative system answering legal questions with no check that its answers matched the law: a disclaimer described the risk, nothing reduced it. An Eval Gate in CI on a legal-question set, a Runtime Guardrail that grounds answers in official sources or refuses, and an Adversarial Red-Team Suite would have caught it before launch. The case touches NIST AI 600-1 on confabulation and, for an EU deployment, AI Act Art. 50(1).

## What happened

In October 2023 New York City announced an AI-powered chatbot to help business owners navigate government; it runs on Microsoft's Azure AI services [1].

The Markup reported on 29 Mar 2024 that in its testing the bot said landlords did not have to accept tenants with housing vouchers, although source-of-income discrimination is illegal in the city, and told a user they could take a cut of workers' tips. On the voucher question the bot once told a reporter that landlords did have to accept vouchers, then told ten separate staffers that they did not [1].

A city spokesperson said the chatbot was a pilot that would improve, had already given thousands of people accurate answers, and disclosed its risks to users [1]. The OECD AI Incidents Monitor and the AI Incident Database both record the case [2][3].

## Failure mode

A generative system answered legal questions in a domain where a wrong answer invites unlawful conduct, with no check that answers matched the law, and with answers that changed from one asking to the next. A disclaimer told users about the risk; nothing reduced it.

## Which control would have caught it

An eval gate built on a legal-question set written with the agencies that enforce each rule (tips, housing, cash acceptance) measures the error rate before launch and sets a bar to clear. A runtime guardrail that grounds answers in official sources and refuses when it cannot cite one, and a red-team pass over the questions the public will ask, close the gap between a pilot label and a public service.

Patterns: [Eval Gate in CI](https://aigovernanceengineer.com/bok/patterns#pattern-eval-gate-in-ci) · [Runtime Guardrail](https://aigovernanceengineer.com/bok/patterns#pattern-runtime-guardrail) · [Adversarial Red-Team Suite](https://aigovernanceengineer.com/bok/patterns#pattern-adversarial-red-team-suite)

## The evidence that would have existed

What an auditor could have read, and the stack layer that produces it.

- Layer 3 (Evals & Red Teaming as Evidence): Eval report on the legal-question set with the launch threshold and error rate per topic
- Layer 3 (Evals & Red Teaming as Evidence): Consistency results: the same question asked many times, with the spread of answers
- Layer 4 (Runtime Controls & Observability): Guardrail logs with the official source cited for each answer, or the refusal
- Layer 3 (Evals & Red Teaming as Evidence): Red-team findings, each closed or formally accepted before launch

## Obligations it touches today

As of 2026-09-24. Mappings are illustrative, not a claim of conformity.

- NIST AI 600-1 [Confabulation](https://aigovernanceengineer.com/obligations/aige-obl-nist-ai600-1): The generative-AI profile lists confabulation, confidently stated but false content, among its twelve risks [4].
- EU AI Act [Art. 50(1)](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art50): For an EU deployment, people must be told they are dealing with an AI system [5]. The duty is about disclosure, not accuracy, which is why the eval gate matters more than the banner.

## System boundary

The MyCity chatbot as the city launched it: a generative model on Microsoft's Azure AI services, the questions business owners typed and the answers it returned about city rules on housing, employment and running a business [1]. The agencies' rules sit outside the system as its reference; the people who act on an answer sit outside it too, and that is where the harm lands.

## Control assumptions

What the controls below take for granted. Challenge any of them.

- An answer about the law is correct only if it matches the rule the enforcing agency applies; fluency and confidence are not evidence of either.
- The same question can get different answers: in The Markup's testing the voucher answer changed from one asking to the next [1], so one passing run of a test set shows little.
- A disclaimer describes the risk without reducing it; the error rate falls only through a check before launch and a guardrail at the point of output.

## Controls by moment

- Preventive: [Eval Gate in CI](https://aigovernanceengineer.com/bok/patterns#pattern-eval-gate-in-ci) · [Adversarial Red-Team Suite](https://aigovernanceengineer.com/bok/patterns#pattern-adversarial-red-team-suite)
- Detective: [Runtime Guardrail](https://aigovernanceengineer.com/bok/patterns#pattern-runtime-guardrail) · [Continuous Assurance Telemetry](https://aigovernanceengineer.com/bok/patterns#pattern-continuous-assurance-telemetry)
- Responsive: [Incident Pipeline](https://aigovernanceengineer.com/bok/patterns#pattern-incident-pipeline) · [Kill Switch / Circuit Breaker](https://aigovernanceengineer.com/bok/patterns#pattern-kill-switch--circuit-breaker)

## Evidence requirements

The evidence each control must leave, written as acceptance criteria.

- An eval report on a versioned legal-question set, written with the agencies that enforce each rule, shows the error rate per topic below the launch threshold for the version that goes live.
- Consistency results show each question in the set asked many times, with the spread of answers recorded and within the threshold: a validity check that the eval measured the answers people get, not one lucky run.
- Every answer in the guardrail log cites the official source it rests on, or is a refusal.
- Every red-team finding is closed, or formally accepted with an owner, before launch.

## Related open controls

Draft control specifications from the open control profiles, open for technical review.

- [AIGE-CTL-EVAL-009](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-009) Evaluation Validity Checks

## Open questions

- What error rate on a legal-question set is low enough to launch a public service that answers questions about the law, and who signs that threshold off?
- How should an eval gate score a question that is answered correctly on one asking and wrongly on the next?

## How to read this case

Each case is an illustrative engineering analysis of public records, not a legal determination, not a finding of fact beyond what the cited sources state, and not a claim of conformity. Mappings to obligations are illustrative.

## Sources

[1] NYC's AI Chatbot Tells Businesses to Break the Law (investigative testing of the MyCity chatbot). The Markup. 2024-03-29. https://themarkup.org/news/2024/03/29/nycs-ai-chatbot-tells-businesses-to-break-the-law (verified: secondary)
[2] OECD AI Incidents Monitor: NYC MyCity Chatbot Gives Dangerous, Illegal Advice to Businesses. OECD.AI. 2024-03-29. https://oecd.ai/en/incidents/2024-03-29-3dce (verified: primary)
[3] AI Incident Database, Incident 714: Microsoft-Powered New York City Chatbot Advises Illegal Practices. Responsible AI Collaborative. 2026. https://incidentdatabase.ai/cite/714/ (verified: primary)
[4] Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1; confabulation listed among twelve generative-AI risks). NIST. 2024-07-26. https://doi.org/10.6028/NIST.AI.600-1 (verified: primary)
[5] EU AI Act Art. 50 (transparency; systems that interact directly with natural persons must be designed so that those persons are informed they are interacting with an AI system). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_50 (verified: primary)
