How much can six months tell us? NYC comparator methods lab

Demonstrating robustness checks for cityoutcomes.com with synthetic data

A synthetic-data demonstration of baseline, weighting, placebo, and trend-sensitivity checks for NYC homicide comparisons.
Author

Kremu

Published

October 1, 2026

Show code
({
  (dplyr)
  (ggplot2)
})
("nyc-comparator/R/rebuttal_funs.R")
("nyc-comparator/R/simulate_demo.R")
("nyc-comparator/R/verify_dashboard.R")

DEMO <- !(params$data_path)
raw  <- if (DEMO) simulate_demo() else (params$data_path)
d    <- prep_panel(raw) |> balance_panel(params$start, params$end)
verify_dashboard_inputs("nyc-comparator", DEMO, params$start, params$end, params$event, params$focus)

focus <- params$focus
event <- (params$event)
pre <- params$pre; n_post <- params$n_post
(focus %in% d$city)

g_eq  <- gap_series(d, focus, "equal")
g_pop <- gap_series(d, focus, "pop")

nyc_base_rate <- d |>
  (city == focus, month < event) |>
  (n = pre) |>
  ((rate)) |>
  ()
pct <- function(x) ("%+.0f%%", 100 * x / nyc_base_rate)
pm <- function(x) ("±%.0f%%", (100 * x / nyc_base_rate))
f2  <- function(x) ("%.2f", x)
((base_size = 11))

DEMO DATA. No input file was found at the configured path, so every number below comes from a synthetic panel. These simulated numbers cannot support claims about actual NYC outcomes.

The bundled simulation uses seed 2026 and sets NYC’s added post-event effect to zero. Shared shocks multiply the cities’ different baseline rates, so additive level-gap parallel trends need not hold even without a policy effect. Thus, this demo combines count noise with simulated drift, seasonality and model dependence. It does not measure power to detect an observed policy effect.

Six months is not a lot of data. The compares New York with a set of peer cities, and its homicide tab (checked October 1, 2026) discloses limited pretrend power and the components of its uncertainty. Here, we’d like to walk through how much the answer from a design like that can move when we change the baseline, the peer weights, or the noise model.

This methods lab adapts material supplied by Kremu. All bundled results use simulated data, not observed NYC outcomes. Thus, it evaluates analytical choices, and it says nothing about the effects of Zohran Mamdani’s policies. We would need real data, and a reconciliation with that tab, before calling any of this an empirical replication or refutation.

There is also an that runs the same estimators. If some of the terms below are unfamiliar, the posts on and cover the background.

What this document does and does not claim

It does not claim the new administration changed homicide in either direction.

What it does is show how baseline choices, weighting, placebo comparisons, and assumed drift change an analysis. However, with synthetic inputs, it cannot test the site’s empirical conclusion.

Fixed primary specification

These are the analysis choices that came with the supplied material. No timestamped registration record was included, so we don’t claim preregistration here. Still, fixing one specification up front helps us tell the primary calculation apart from the exploratory settings.

Choice Value Why
Outcome Monthly homicide rate per 100k The site’s headline tab
Event date Jan 2026 Same as the site
Post window 6 months All available data
Primary baseline Mean of the last 12 pre-event months A single-month baseline is noisy
Replication baseline t−1 only, equal weights Matches the conventional baseline and weighting rule; real-data replication remains unverified
Peer weights Equal (primary) and population-pooled The site uses equal
Diagnostics Placebo ranks, three noise summaries and a historical-window sensitivity One treated unit
Sensitivity Year-over-year seasonal check; linear-drift tipping point Pretrend power is low

1. Demonstrating the conventional estimator

Show code
site_est <- did_estimate(g_eq, event, pre, n_post, baseline = "last")
ep <- event_path(g_eq, event) |> (rel_month >= -48)
latest <- (ep$effect, 1)

Let’s start with the conventional estimator. Using the site’s method (equal-weighted peers, rebased to t−1), the mean post-period change in the NYC–peer gap relative to December 2025 is 0.03 per 100k/month. The latest single-month gap change is 0.07.

Show code
(ep, (rel_month, effect)) +
  (yintercept = 0, colour = "grey60") +
  (xintercept = -0.5, linetype = 2) +
  () + (size = 0.8) +
  (x = "Months relative to event", y = "NYC minus peers (per 100k)")
Synthetic focal-city minus peer-rate changes relative to December 2025.
Figure 1: Synthetic event-study path (rebased so t−1 = 0).

2. The point estimate depends on unforced choices

Show code
variants <- tibble::(
  Specification = ("Site: t-1 baseline, equal weights",
                    "12-month mean baseline, equal  (primary)",
                    "12-month mean baseline, population weights",
                    "6-month mean baseline, equal weights",
                    "24-month mean baseline, equal weights",
                    "Year-over-year, same  (seasonality-robust)"),
  Estimate = (site_est,
               did_estimate(g_eq,  event, 12, n_post),
               did_estimate(g_pop, event, 12, n_post),
               did_estimate(g_eq,  event, 6,  n_post),
               did_estimate(g_eq,  event, 24, n_post),
               yoy_estimate(g_eq, event, n_post))
) |> (`% of NYC baseline` = pct(Estimate))
knitr::(variants |> (Estimate = f2(Estimate)))
Specification Estimate % of NYC baseline
Site: t-1 baseline, equal weights 0.03 +9%
12-month mean baseline, equal weights (primary) 0.00 +1%
12-month mean baseline, population weights -0.06 -18%
6-month mean baseline, equal weights 0.06 +18%
24-month mean baseline, equal weights 0.02 +6%
Year-over-year, same months (seasonality-robust) -0.05 -16%
Show code
primary <- did_estimate(g_eq, event, pre, n_post)

NYC’s baseline rate is 0.32 per 100k per month. As we can see from the table, the estimate depends on choices that nothing forces on us. That spread across defensible specifications is itself a form of uncertainty, and the illustrative intervals below condition on whichever specification was chosen.

A note on the percentage column. It divides the estimated change in the NYC–peer gap by NYC’s 12-month baseline rate, so it should not be read as NYC’s own percentage rate change (both NYC and peer rates enter the contrast). Thus, a value below −100% can occur even when every homicide rate is nonnegative.

3. How big an effect could six months detect?

With a single treated unit, the variance model does most of the work. The site’s methodology, checked October 1, 2026, discloses cross-peer variation, Poisson count variance, and Newey–West slope uncertainty. We don’t reproduce its full interval calculation here. Instead, here’s a table of three illustrative noise summaries and a historical-window sensitivity:

Show code
fake <- (("2019-01-01"), ("2025-01-01"), by = "year")
itp  <- in_time_placebos(g_eq, fake, pre, n_post)
isp  <- in_space_placebos(d, event, pre, n_post, params$noise_window)
n_itp <- (!(itp$estimate))

se_tbl <- tibble::(
  `Variance model` = ("Focal count-only component (independent Poisson reference)",
                       "NYC's own in-time placebos",
                       "Cross-city placebo  (alternative diagnostic)",
                       "Own history omitting 2020–2023 events (sensitivity)"),
  SE = (poisson_se(d, focus, event, pre, n_post),
         (itp$estimate, na.rm = TRUE),
         (isp$estimate[isp$city != focus], na.rm = TRUE),
         (itp$estimate[(itp$fake_event, "%Y") %in% ("2019", "2024", "2025")], na.rm = TRUE))
) |>
  (`Normal reference band (±1.96 × summary)` = ("%s to %s", f2(primary - 1.96 * SE), f2(primary + 1.96 * SE)),
         `MDE (80% power)` = f2(mde(SE)),
         `MDE as % of NYC baseline` = (mde(SE)))
knitr::(se_tbl |> (SE = f2(SE)))
Variance model SE Normal reference band (±1.96 × summary) MDE (80% power) MDE as % of NYC baseline
Focal count-only component (independent Poisson reference) 0.03 -0.05 to 0.06 0.08 ±26%
NYC’s own in-time placebos 0.12 -0.23 to 0.23 0.33 ±103%
Cross-city placebo spread (alternative diagnostic) 0.28 -0.55 to 0.56 0.79 ±249%
Own history omitting 2020–2023 events (sensitivity) 0.03 -0.05 to 0.06 0.08 ±26%

The minimum detectable effect (MDE) shown here is a normal-approximation planning calculation for 80% power at a two-sided 5% threshold, conditional on the chosen standard error and model. Smaller effects would be detected less often. However, these approximate summaries do not establish the actual power of the site’s design.

These summaries measure different sources of variation. The count-only component excludes peer counts and dependence. The in-time spread includes count variation and historical change, while the cross-city spread also includes differences between cities. Thus, we treat the bands as descriptive normal references (±1.96 times each summary), without demonstrated 95% coverage. The sensitivity row uses only the 2019, 2024 and 2025 events. It exposes the consequence of omitting 2020–2023 events, but does not establish that three draws provide a better reference. Choosing that summary also changes the approximate MDE and compatibility curve.

Differences in city size can create heteroskedasticity and distort inference with few treated groups. That said, this synthetic example does not establish that the site’s intervals overstate NYC uncertainty. The in-time placebos use NYC’s own history, which seems like the most relevant reference, but there are only 7 draws, so we should treat that SE as rough. The intervals above also use the normal multiplier 1.96. An illustrative t reference with 6 degrees of freedom and the same SD would use 2.45 instead, which widens the in-time interval by about 25%. The placebo windows overlap from one year to the next as well, so the draws are not independent, and the t reference should not be taken as a validated correction.

4. Which effect sizes are compatible with the data?

Show code
se_mid <- se_tbl$SE[2]
hyp <- nyc_base_rate * (-0.5, -0.25, 0, 0.25, 0.5)
sv <- s_values(primary, se_mid, hyp) |>
  (`Hypothesized change` = pct(hypothesis),
         p = ("%.2f", p), `S- (bits)` = ("%.1f", s_bits)) |>
  (`Hypothesized change`, p, `S- (bits)`)
knitr::(sv)
Hypothesized change p S-value (bits)
-50% 0.16 2.6
-25% 0.48 1.1
+0% 0.98 0.0
+25% 0.51 1.0
+50% 0.18 2.5

The table uses NYC’s in-time SE. An S-value of s bits is about as surprising as seeing s heads in a row from a fair coin (the goes through this in detail). An S-value below 2 bits corresponds to a P-value above 0.25 under the stated normal model. It is NOT the probability that a hypothesis is true.

Show code
grid <- (primary - 4 * (se_tbl$SE), primary + 4 * (se_tbl$SE), length.out = 400)
curves <- (((se_tbl)), function(i)
  (model = se_tbl$`Variance model`[i], h = grid,
             p = s_values(primary, se_tbl$SE[i], grid)$p)) |> ()
(curves, (h, p, linetype = model)) +
  () + (xintercept = 0, colour = "grey60") +
  (x = "Hypothesized effect (per 100k/month)", y = "p-value", linetype = NULL) +
  (legend.position = "bottom", legend.direction = "vertical")
Compatibility curves under four illustrative noise summaries, including a historical-window sensitivity.
Figure 2: P-value (compatibility) function for the primary estimate under each variance model.

5. Is NYC unusual among its peers? (in-space permutation)

Show code
nyc_row  <- isp |> (city == focus)
p_ratio  <- perm_p(isp$ratio, nyc_row$ratio)
p_est    <- perm_p(isp$estimate / isp$pre_rmspe, nyc_row$estimate / nyc_row$pre_rmspe)
rank_r   <- (-(isp$ratio, !(isp$ratio), NA_real_), ties.method = "min", na.last = "keep")[isp$city == focus]

Each city takes a turn as the focus against all other cities, so NYC remains in a peer city’s comparator pool. Both RMSPE windows use that city’s pre-baseline mean as their center, with a 24-month pre noise window. We retain this descriptive leave-one-out diagnostic, but a real treatment effect could also enter placebo comparator pools.

Among 30 finite ratios, NYC ranks 16. The upper-tail ratio share is 0.53, and the two-sided share for the standardized estimate is 0.97. The ratio share has a floor of 1/30 when the focal ratio is finite and included. These are descriptive rankings, which are a long way from a fitted synthetic-control or calibrated randomization analysis.

Show code
(isp, ((city, ratio), ratio, fill = city == focus)) +
  () + () +
  (values = ("grey70", "firebrick"), guide = "none") +
  (x = NULL, y = "Post / pre RMSPE")
Synthetic cities ranked by post-period to pre-period RMSPE ratio.
Figure 3: Post/pre RMSPE ratio for every city treated as the focus.

6. Is January 2026 unusual in NYC’s own history? (in-time placebos)

Show code
itp_plot <- (itp |> (type = "Placebo"),
                      tibble::(fake_event = event, estimate = primary, type = "Actual"))
(itp_plot, (fake_event, estimate, colour = type)) +
  (yintercept = 0, colour = "grey60") +
  (size = 2.5) +
  (values = (Actual = "firebrick", Placebo = "grey40")) +
  (x = "Event date", y = "DiD estimate (per 100k/month)", colour = NULL)
p_time <- perm_p((itp$estimate, primary), primary)
Synthetic estimates at January placebo dates and the analysis date.
Figure 4: NYC estimate at fake January event dates vs the real one.

Across 8 January start dates (the analysis date included), the share with an estimate at least as large in magnitude as the analysis-date estimate is 1.00. Missing-window estimates are excluded from both the share and its denominator. However, the 2020–2023 event windows include simulated pandemic-era shocks, so the sensitivity row above shows the spread after omitting them. The 2025 placebo’s post window also overlaps the primary baseline, making these draws dependent.

7. How much trend violation would flip the reading?

Show code
sl  <- pre_slope(g_eq, event, params$noise_window)
lag <- trend_lag(pre, n_post)
delta_grid <- (-3, 3, length.out = 121) * ((sl["slope"]), sl["se"])
adj <- drift_adjust(primary, se_mid, delta_grid, lag)
breakdown <- primary / lag

The pre-period slope of the gap is 0.0002 per month (naive OLS SE 0.0038; its ±1.96 × SE reference runs from -0.0073 to 0.0077), and the baseline and post midpoints are 9 months apart. This reference ignores serial correlation and has no demonstrated coverage. Thus, a differential drift of just 0.0004 per month would fully account for the point estimate.

Show code
(adj, (delta, estimate)) +
  ("rect", xmin = sl["slope"] - 1.96 * sl["se"], xmax = sl["slope"] + 1.96 * sl["se"],
           ymin = -Inf, ymax = Inf, alpha = 0.15) +
  ((ymin = lo, ymax = hi), alpha = 0.25) +
  () + (yintercept = 0, linetype = 2) +
  (x = "Assumed  (per 100k per month)", y = "Adjusted estimate")
Adjusted synthetic estimate and interval across assumed linear drifts.
Figure 5: Estimate after removing an assumed linear differential drift. Shaded: naive OLS slope reference (±1.96 SE), ignoring serial correlation.

This is a linear drift scenario, motivated by work on sensitivity to parallel-trend violations. It does not implement HonestDiD, and it does not propagate uncertainty in the assumed drift.

8. Scorecard

Show code
tibble::(
  Check = ("Conventional estimator", "Specification  (six report choices)", "Approximate MDE (in-time summary)",
            "Permutation rank", "In-time rank", "Drift breakdown"),
  Result = (f2(site_est),
             ("%s to %s", f2((variants$Estimate, na.rm = TRUE)),
                     f2((variants$Estimate, na.rm = TRUE))),
             (mde(se_mid)),
             ("%d / %d (tail share = %.2f)", rank_r, ((isp$ratio)), p_ratio),
             ("p = %.2f", p_time),
             ("%.4f per month", breakdown))
) |> knitr::()
Check Result
Conventional estimator 0.03
Specification spread (six report choices) -0.06 to 0.06
Approximate MDE (in-time summary) ±103%
Permutation rank 16 / 30 (tail share = 0.53)
In-time rank p = 1.00
Drift breakdown 0.0004 per month

Limitations of this methods demonstration

  • We were not given a real homicide input file or a preregistration record. Thus, no empirical replication or policy-effect conclusion has been established.
  • Placebo ranks require exchangeability to be interpreted as calibrated P-values. Cities and historical dates are not randomly assigned, so we treat these ranks as descriptive diagnostics. The normal intervals and MDEs condition on illustrative variance models.
  • We only cover homicide. Other outcomes would need separate checks of their measurement, geography, timing and comparator definitions.
  • Council on Criminal Justice (CCJ) mid-year data are preliminary and subject to revision, so it’s worth recording the date the data were pulled.
  • The permutation p has a floor of 1/N. The in-time placebos have few draws.
  • The linear-drift calculation is illustrative and does not implement HonestDiD robust inference.

As a rule, we’d look to the whole specification table and the interval estimates before any single point estimate, and we’d hold off on any reading of the real site until the real data have been run through the same checks.

References

1. Ferman B, Pinto C. (2019). “Inference in differences-in-differences with few treated groups and heteroskedasticity.” The Review of Economics and Statistics. 101:452–467. doi: .
2. Rambachan A, Roth J. (2023). “A more credible approach to parallel trends.” The Review of Economic Studies. 90:2555–2591. doi: .
Back to top

Comments

Webmentions

Replies, likes, and mentions from around the web, tracked via .