Background

Last week’s assignment looked at three individual readings from this same noise sensor network, the loudest, the quietest, and the one closest to average, and found that one of those three, a 0.0 dB reading, might not have even been real data. This week builds directly on that, but shifts the focus from single rows to patterns across the whole dataset, specifically time, geography, and temperature, since those are dimensions this sensor network is well suited to capture.

The idea of a city having a “pulse,” shifts in conditions and behavior that make up daily urban life, is normally hard to see directly, since the datasets and visuals I’ve seen usually flatten it into monthly or yearly averages. But because these sensors report roughly once a minute, that pulse becomes much more visible. This assignment tries to expose some of that pulse, in time, in geography, weather, etc to think about what a large dataset like this can, and can’t, actually tell you about the surrounding areas.

Analysis

The Data and Classifying the Variables

library(tidyverse)

# read in the sensor readings for May 2026
sensor_data <- read_csv("nu_readings_2026_05.csv")

# look at the type of each variable in the dataset
str(sensor_data)
## spc_tbl_ [1,710,639 × 8] (S3: spec_tbl_df/tbl_df/tbl/data.frame)
##  $ reading_id   : num [1:1710639] 12947174 12947199 12947233 12947261 12947286 ...
##  $ deployment_id: num [1:1710639] 11 13 14 15 16 17 18 1 60 20 ...
##  $ location_id  : num [1:1710639] 11 13 10 15 16 17 18 1 8 20 ...
##  $ timestamp    : POSIXct[1:1710639], format: "2026-05-01 00:00:00" "2026-05-01 00:00:00" ...
##  $ temperature  : num [1:1710639] 54.2 55.9 55.5 54.9 55.3 54.3 55.4 54.4 54.7 54.6 ...
##  $ humidity     : num [1:1710639] 88.5 100 94.1 100 100 100 100 100 100 88.4 ...
##  $ noise        : num [1:1710639] 44.4 59.3 52.3 53.7 69.2 45.6 60.7 66.1 53.7 48.3 ...
##  $ heat_index   : num [1:1710639] 53.8 55.9 55.3 54.8 55.3 ...
##  - attr(*, "spec")=
##   .. cols(
##   ..   reading_id = col_double(),
##   ..   deployment_id = col_double(),
##   ..   location_id = col_double(),
##   ..   timestamp = col_datetime(format = ""),
##   ..   temperature = col_double(),
##   ..   humidity = col_double(),
##   ..   noise = col_double(),
##   ..   heat_index = col_double()
##   .. )
##  - attr(*, "problems")=<pointer: 0xa69f802b0>

The dataset has 8 variables. The reading_id is just a unique ID number for each row, it doesn’t really describe anything about the world. deployment_id and location_id are categorical, they identify which sensor unit and which physical spot took the reading. timestamp is a date-time variable, which can be broken down into other categorical pieces like hour of day or day of week. temperature, humidity, noise, and heat_index are all continuous numeric variables, actual measurements that reflect a wide range of values.

Pattern 1: The Daily Rhythm (Time)

# average noise by hour of day, across every sensor
sensor_data %>%
  mutate(hour_of_day = hour(timestamp)) %>%
  group_by(hour_of_day) %>%
  summarize(avg_noise = mean(noise)) %>%
  ggplot(aes(x = hour_of_day, y = avg_noise)) +
  geom_line() +
  labs(title = "Average Noise by Hour of Day", x = "Hour of Day", y = "Average Noise (dB)")

This is basically the city’s pulse showing up directly in the data. Noise bottoms out around 7am at about 50 dB, then climbs steadily through the day to a peak around 59-60 dB in the early evening, before dropping back off overnight. This shows a real, repeating daily rhythm, not something you’d ever see in a single monthly average like many other datasets, and reflects the power of the granularity of this data.

Pattern 2: Geography

# average noise by location, across the whole month
sensor_data %>%
  group_by(location_id) %>%
  summarize(avg_noise = mean(noise)) %>%
  mutate(location_id = as.factor(location_id)) %>%
  ggplot(aes(x = location_id, y = avg_noise)) +
  geom_col() +
  labs(title = "Average Noise by Sensor Location", x = "Location ID", y = "Average Noise (dB)")

Geography matters just as much as time, maybe more. Average noise by location ranges from about 43 dB all the way up to about 93 dB, basically a 50 dB spread just depending on where in the neighborhood the sensor sits. Most locations cluster somewhere in the middle, but a few sit way above everyone else, the same handful of consistently loud sensors that showed up as that second bump in last week’s histogram.

Turns out Sensor 21, the one tied to the “around 400 planes” community story from last week, actually has the lowest average noise of all 45 sensors. Sensors 13 and 19, tied to the Quincy St and Blue Hill Ave complaint, rank 8th and 15th loudest out of 45, genuinely in the upper half. So one complaint lines up cleanly with the geographic average, and the other one doesn’t line up at all, which says something important about what an average can and can’t capture, more on that below.

One sensor really stands out on the loud end though. Sensor 56 averages 92.7 dB, way ahead of the next loudest sensor at 73.3 dB, almost a 20 dB gap between 1st and 2nd place. And it’s not just the average that’s extreme, Sensor 56’s quietest reading all month was 89.6 dB, louder than the average of every other sensor in the entire network. This isn’t a location with occasional loud spikes the way Sensor 21 is, it’s basically never quiet, which points to a completely different kind of noise problem than anything else in this dataset, probably something running constantly right next to it, like the HVAC example from class.

Before going further with this, it’s worth checking whether Sensor 56 even has enough data to trust, maybe it’s just a handful of readings skewing the average, rather than a real pattern.

# total readings for sensor 56 specifically
sensor_data %>%
  filter(location_id == 56) %>%
  nrow()
## [1] 38484
# compare that to the typical (median) number of readings across all 45 sensors
sensor_data %>%
  group_by(location_id) %>%
  summarize(n_readings = n()) %>%
  summarize(median_readings = median(n_readings))

Turns out that’s not the issue. Sensor 56 has 38,484 readings for the month, right in line with the network median of about 38,987. So this isn’t a case of a handful of readings skewing the average, it’s a full month’s worth of data, basically the same amount as almost every other sensor. If anything, that makes the finding more convincing, not less, it’s not a fluke from a small sample, it’s a real, sustained pattern held up across a normal amount of data.

I tried checking the Common SENSES live map to see if there was anything obvious near Sensor 56 that could actually explain this, a highway, a construction site, something running constantly nearby, but I couldn’t actually find Sensor 56 on the map at all. So for now this one’s just a striking number without a story behind it, I don’t have a good explanation for why it’s this loud, just that it clearly, consistently is.

Per the feedback, I also went looking for a separate location dataset that might have an address attached to each sensor, since that would be the more direct way to check this than point-and-click on the map. I wasn’t able to find one though, only the main readings file, which has no address information in it either.

Since I can’t explain it, it’s worth checking how much Sensor 56 is actually driving the geography pattern above, versus how much of that spread is real across the rest of the network.

# redo the geography comparison excluding sensor 56
sensor_data %>%
  filter(location_id != 56) %>%
  group_by(location_id) %>%
  summarize(avg_noise = mean(noise)) %>%
  summarize(spread = max(avg_noise) - min(avg_noise))

Turns out Sensor 56 is doing a lot of the work here. With it included, the gap between the loudest and quietest sensor is about 49 dB. Take it out, and that gap drops to about 30 dB, nearly cut in half. So the geography story is still real, there’s still a genuine 30 dB difference across the rest of the network, but a good chunk of how dramatic that spread looked was really just one unexplained sensor, not a broad pattern across the whole neighborhood.

Pattern 3: Temperature and Noise

# check the correlation between temperature and noise
cor(sensor_data$temperature, sensor_data$noise)
## [1] 0.1502714

Temperature and noise turn out to be related too, though only loosely, the correlation’s about 0.15, which isn’t strong on its own. Breaking readings into 10-degree temperature ranges shows a cleaner trend though:

# average noise across 10-degree temperature ranges
sensor_data %>%
  mutate(temp_bucket = cut(temperature, breaks = seq(30, 100, by = 10))) %>%
  group_by(temp_bucket) %>%
  summarize(avg_noise = mean(noise)) %>%
  ggplot(aes(x = temp_bucket, y = avg_noise)) +
  geom_col() +
  labs(title = "Average Noise by Temperature Range", x = "Temperature (°F)", y = "Average Noise (dB)")

Average noise climbs steadily from about 53 dB in the 40s up to about 60.5 dB in the 90s, over 7 dB warmer-to-hotter. It’s not a huge effect, but it’s a steady, one-directional one, not just noise in the correlation. Makes some sense too, warmer weather probably means more people outside, more open windows instead of closed ones, and more AC units running, all things that would nudge ambient noise up a bit as it gets hotter.

Based on the feedback on this section, I decided to explore the noise-temperature relationship a bit deeper rather than leave it at speculation. That explanation makes a testable prediction: if it’s really about people being outside and windows being open, the effect should be weaker overnight when most people are asleep, regardless of temperature. So I checked that directly.

# check the temperature-noise correlation separately for daytime vs overnight hours
sensor_data %>%
  mutate(hour_of_day = hour(timestamp)) %>%
  filter(hour_of_day >= 8, hour_of_day < 22) %>%
  summarize(daytime_correlation = cor(temperature, noise))
sensor_data %>%
  mutate(hour_of_day = hour(timestamp)) %>%
  filter(hour_of_day < 8 | hour_of_day >= 22) %>%
  summarize(overnight_correlation = cor(temperature, noise))

Turns out the opposite of what I expected. The correlation is actually a little stronger overnight (about 0.14) than during the day (about 0.12). If this were really about people being outside more or opening windows when it’s warm, the effect should have been weaker overnight, not stronger. That pushes me away from the “more people around” explanation and more toward something mechanical, like air conditioning units, which might run just as much, or even more, overnight to cool a home down after a hot day, regardless of whether anyone’s outside.

Community Stories

Story about noise (sensors 13, 19), near Quincy St and Blue Hill Ave: “Noise pollution, airplanes, screeching tires, sirens. Suffering on side street because drivers are avoiding main street.”

Second story about noise (sensor 21), near a home under a flight path: “Planes are all routed through this area, extremely loud. Around 400 planes flying over house.”

These two stories map onto the geography pattern in two completely different ways. The Quincy St and Blue Hill Ave complaint lines up cleanly with the data, Sensors 13 and 19 really are among the loudest locations in the network on average, so the resident’s experience and the aggregate statistic agree. The flight path complaint doesn’t line up at all, Sensor 21 has the lowest average of any sensor in the whole network, yet the resident describes it as extremely loud. That’s not a contradiction so much as a limit of what an average can show, a location can experience real, disruptive events and still look calm in a geographic ranking, if those events are rare enough to get buried by everything quiet in between.

Interpretation

Each row in this dataset is simply one place in time with a couple of measurements. On its own, that tells you almost nothing about the neighborhood. But once you aggregate across 1.7 million of them, patterns emerge that no individual reading could show, a repeating daily rhythm, a real geographic divide between loud and quiet parts of the neighborhood, even a loose link to the weather. That’s basically the tradeoff, you get to see the big population-level patterns, but only by flattening out what any one person might actually experience. Sensor 21 is the clearest example, at the population level, one number for the whole sensor, it disappears into the quiet end of the distribution. At the individual level, for the person living under that flight path, as shown through the community story, it could be one of the worst spots in the entire network.

That daily rhythm is basically what “the pulse of the city” means, quiet in the early morning, loud in the evening, a shift that lines up with when people are asleep, waking up, commuting, and out and about. That pulse has always existed, but it used to be much harder to see, you’d need someone standing on every corner with a decibel meter around the clock to catch it. A sensor network that reports every minute makes that rhythm directly visible. Geography adds a second layer on top of it, the whole neighborhood doesn’t just breathe in the same rhythm together, different parts of it breathe at completely different volumes. Temperature adds a third, quieter layer still.

Taken together, this shows that time, geography, and temperature are all doing real work here, and that averages can hide pretty different things depending on which one you’re looking at. A single sensor’s monthly average can hide a rhythmic daily pattern entirely. A neighborhood-wide geographic average can hide the fact that one specific location has a real problem, if that problem happens rarely enough.

This raises further questions worth digging into. If Sensor 21’s average is this misleading, are there other quiet-on-average sensors hiding a similar pattern, real but infrequent disruptive events that would only show up if someone specifically went looking for spikes rather than trusting the average? Last week’s analysis found one example of that, and this week’s geographic view suggests there could be more. Sensor 56 raises a related but different question: since I couldn’t locate it on the Common SENSES map, and the dataset itself has no address information either, this is worth following up on directly, maybe by reaching out to whoever maintains the sensor network, rather than something I can fully resolve from the data alone.

AI use

I used Claude to help debug the ggplot code for my visualizations this week, to help think through how to connect the time, geography, and temperature patterns back to the community stories and the “pulse of the city” idea from the module overview, to help identify Sensor 56 as a standout worth flagging even though I wasn’t able to find a real-world explanation for it, to help check whether Sensor 56’s high readings could be explained by a small sample size, which turned out not to be the case, and to help design and run the daytime-versus-overnight test on the temperature-noise relationship. I also used Claude for {r, fig.width=10, fig.height=5} as the graph was too crammed otherwise.