Texas Realty Insights

This project analyses historical real estate market data for four Texas cities between 2010 and 2014.

The objective is to identify sales trends, evaluate variability and distributional characteristics, investigate seasonal patterns, compare cities and periods, and provide statistical insights that may support business decisions.

case <- read.csv("realestate_texas.csv")

library(ggplot2)

str(case)
## 'data.frame':    240 obs. of  8 variables:
##  $ city            : chr  "Beaumont" "Beaumont" "Beaumont" "Beaumont" ...
##  $ year            : int  2010 2010 2010 2010 2010 2010 2010 2010 2010 2010 ...
##  $ month           : int  1 2 3 4 5 6 7 8 9 10 ...
##  $ sales           : int  83 108 182 200 202 189 164 174 124 150 ...
##  $ volume          : num  14.2 17.7 28.7 26.8 28.8 ...
##  $ median_price    : num  163800 138200 122400 123200 123100 ...
##  $ listings        : int  1533 1586 1689 1708 1771 1803 1857 1830 1829 1779 ...
##  $ months_inventory: num  9.5 10 10.6 10.6 10.9 11.1 11.7 11.6 11.7 11.5 ...

Variable Analysis

The dataset contains 240 observations covering four cities over 60 months, from January 2010 to December 2014.

Each row represents one city in one specific month and year. The dataset therefore combines a geographical dimension with a time dimension.

The variables can be classified as follows:

Position, variability and shape measures are appropriate for the quantitative variables sales, volume, median_price, listings and months_inventory.

To correctly analyse the chronological evolution of the market, year and month are combined into a single date variable.

case$date <- as.Date(
  sprintf("%04d-%02d-01", case$year, case$month)
)

colSums(is.na(case))
##             city             year            month            sales 
##                0                0                0                0 
##           volume     median_price         listings months_inventory 
##                0                0                0                0 
##             date 
##                0
sum(duplicated(case[c("city", "year", "month")]))
## [1] 0

There are no missing values and no duplicated combinations of city, year and month.

Measures of Position, Variability and Shape

Frequency distributions are appropriate for the categorical and time variables.

table(case$city)
## 
##              Beaumont Bryan-College Station                 Tyler 
##                    60                    60                    60 
##         Wichita Falls 
##                    60
table(case$year)
## 
## 2010 2011 2012 2013 2014 
##   48   48   48   48   48
table(case$month)
## 
##  1  2  3  4  5  6  7  8  9 10 11 12 
## 20 20 20 20 20 20 20 20 20 20 20 20

Each city has 60 observations, each year contains 48 observations and each calendar month contains 20 observations.

This confirms that all four cities are represented throughout the same five-year period.

Sales

summary(case$sales)
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##    79.0   127.0   175.5   192.3   247.0   423.0
sd(case$sales)
## [1] 79.65111
var(case$sales)
## [1] 6344.3
IQR(case$sales)
## [1] 120
sd(case$sales) / mean(case$sales) * 100
## [1] 41.42203

Monthly sales per city have a mean of approximately 192.29 and a median of 175.50.

The standard deviation is approximately 79.65, while the interquartile range is 120.

The mean is higher than the median, which is consistent with a positively skewed distribution.

Other Quantitative Variables

summary(
  case[c(
    "volume",
    "median_price",
    "listings",
    "months_inventory"
  )]
)
##      volume        median_price       listings    months_inventory
##  Min.   : 8.166   Min.   : 73800   Min.   : 743   Min.   : 3.400  
##  1st Qu.:17.660   1st Qu.:117300   1st Qu.:1026   1st Qu.: 7.800  
##  Median :27.062   Median :134500   Median :1618   Median : 8.950  
##  Mean   :31.005   Mean   :132665   Mean   :1738   Mean   : 9.193  
##  3rd Qu.:40.893   3rd Qu.:150050   3rd Qu.:2056   3rd Qu.:10.950  
##  Max.   :83.547   Max.   :180000   Max.   :3296   Max.   :14.900
sapply(
  case[c(
    "volume",
    "median_price",
    "listings",
    "months_inventory"
  )],
  sd
)
##           volume     median_price         listings months_inventory 
##        16.651447     22662.148687       752.707756         2.303669
sapply(
  case[c(
    "volume",
    "median_price",
    "listings",
    "months_inventory"
  )],
  var
)
##           volume     median_price         listings months_inventory 
##     2.772707e+02     5.135730e+08     5.665690e+05     5.306889e+00
sapply(
  case[c(
    "volume",
    "median_price",
    "listings",
    "months_inventory"
  )],
  IQR
)
##           volume     median_price         listings months_inventory 
##          23.2335       32750.0000        1029.5000           3.1500

Monthly sales volume averages approximately $31.01 million, with a median of approximately $27.06 million and a standard deviation of approximately $16.65 million.

The mean of the monthly median prices is approximately $132,665.42, while their median is $134,500. The standard deviation is approximately $22,662.15.

Active listings average approximately 1,738.02, with a median of 1,618.50 and a standard deviation of approximately 752.71.

Months of inventory average approximately 9.19 months, with a median of 8.95 months and a standard deviation of approximately 2.30 months.

These statistics combine differences between cities, seasonal variation and changes over time. For this reason, they should also be interpreted together with the conditional analyses presented later.

Coefficients of Variation

The coefficient of variation allows variables with different units and scales to be compared in relative terms.

cv_values <- sapply(
  case[c(
    "sales",
    "volume",
    "median_price",
    "listings",
    "months_inventory"
  )],
  function(x) sd(x) / mean(x) * 100
)

round(cv_values, 2)
##            sales           volume     median_price         listings 
##            41.42            53.71            17.08            43.31 
## months_inventory 
##            25.06

Skewness

The following calculation measures the asymmetry of each quantitative distribution.

skewness_values <- sapply(
  case[c(
    "sales",
    "volume",
    "median_price",
    "listings",
    "months_inventory"
  )],
  function(x) mean(((x - mean(x)) / sd(x))^3)
)

round(skewness_values, 3)
##            sales           volume     median_price         listings 
##            0.714            0.879           -0.362            0.645 
## months_inventory 
##            0.041

Sales show positive skewness, meaning that relatively high monthly sales observations extend the upper tail of the distribution.

Sales volume also has positive skewness and therefore presents some relatively high-value city-month observations.

Median price has modest negative skewness.

Listings have positive skewness, indicating the presence of some relatively high listing counts.

Months of inventory have skewness close to zero, indicating relatively little overall asymmetry.

Variables with the Greatest Variability and Asymmetry

round(cv_values, 2)
##            sales           volume     median_price         listings 
##            41.42            53.71            17.08            43.31 
## months_inventory 
##            25.06
names(cv_values)[which.max(cv_values)]
## [1] "volume"
round(skewness_values, 3)
##            sales           volume     median_price         listings 
##            0.714            0.879           -0.362            0.645 
## months_inventory 
##            0.041
names(skewness_values)[which.max(abs(skewness_values))]
## [1] "volume"

The variable with the greatest relative variability is volume, with a coefficient of variation of approximately 53.71%.

The coefficient of variation is more appropriate than directly comparing variances or standard deviations because the variables have different units and scales.

Sales volume is also the variable with the greatest absolute skewness, approximately 0.879.

Its positive skewness indicates that some city-month observations have particularly high total sales values.

By comparison, months_inventory has skewness close to zero, while median_price has moderate negative skewness.

Sales Classes, Frequencies and Gini Heterogeneity

The quantitative variable sales is divided into five classes.

case$sales_class <- cut(
  case$sales,
  breaks = c(0, 100, 200, 300, 400, 500),
  labels = c(
    "0-99",
    "100-199",
    "200-299",
    "300-399",
    "400-499"
  ),
  right = FALSE
)

class_freq <- table(case$sales_class)

class_freq
## 
##    0-99 100-199 200-299 300-399 400-499 
##      20     127      67      23       3
class_prop <- class_freq / sum(class_freq)

round(class_prop, 3)
## 
##    0-99 100-199 200-299 300-399 400-499 
##   0.083   0.529   0.279   0.096   0.013

The class frequencies are:

The 100-199 class is the most frequent, containing approximately 52.9% of all observations.

barplot(
  class_freq,
  main = "Frequency Distribution of Sales Classes",
  xlab = "Number of Sales",
  ylab = "Number of Observations",
  col = "steelblue"
)

The Gini heterogeneity index is calculated as follows:

gini_heterogeneity <- 1 - sum(class_prop^2)

round(gini_heterogeneity, 3)
## [1] 0.626

The Gini heterogeneity index is approximately 0.626.

A value of zero would indicate that all observations belong to a single class.

With five classes, the maximum possible value is 0.800, which would occur if all classes contained equal proportions of observations.

The observed value therefore indicates that the observations are distributed across several classes, although more than half are concentrated in the 100-199 sales class.

This measure is an index of heterogeneity between categories and should not be confused with the Gini coefficient commonly used to measure inequality.

Its value also depends on the selected class boundaries.

Probability Analysis

The probability calculations are based on the relative frequency of rows in the dataset.

prob_beaumont <- sum(case$city == "Beaumont") / nrow(case)

prob_july <- sum(case$month == 7) / nrow(case)

prob_december_2012 <- sum(
  case$year == 2012 & case$month == 12
) / nrow(case)

prob_beaumont
## [1] 0.25
prob_july
## [1] 0.08333333
prob_december_2012
## [1] 0.01666667

The probability that a randomly selected row refers to Beaumont is:

60 / 240 = 0.25 = 25%

The probability that a randomly selected row refers to July is:

20 / 240 = 0.0833 = 8.33%

The probability that a randomly selected row refers to December 2012 is:

4 / 240 = 0.0167 = 1.67%

These probabilities refer to the frequency of observations in the dataset and not to the probability of an individual property being sold.

Creation of New Variables

Average Property Price

The dataset contains total sales volume and the number of sales.

Since volume is expressed in millions of dollars, it is multiplied by 1,000,000 before being divided by the number of properties sold.

case$average_price <- case$volume * 1000000 / case$sales

head(
  case[c(
    "volume",
    "sales",
    "average_price"
  )]
)
##   volume sales average_price
## 1 14.162    83      170626.5
## 2 17.690   108      163796.3
## 3 28.701   182      157697.8
## 4 26.819   200      134095.0
## 5 28.833   202      142737.6
## 6 27.219   189      144015.9
summary(case$average_price)
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##   97010  132939  156588  154320  173915  213234

Average property prices across city-month observations range from approximately $97,010 to $213,234.

The median monthly average property price is approximately $156,588, while the unweighted mean across city-month observations is approximately $154,320.

average_price represents the arithmetic mean sale price within each city-month.

This differs from median_price, because the arithmetic mean is more strongly influenced by relatively expensive properties.

The mean of average_price across all rows gives equal weight to each city-month and is therefore not the same as the sales-weighted average price of all individual properties sold.

Sales-to-Listings Ratio

A second variable is created to measure monthly sales relative to the number of active listings.

case$sales_to_listings_pct <-
  case$sales / case$listings * 100

head(
  case[c(
    "sales",
    "listings",
    "sales_to_listings_pct"
  )]
)
##   sales listings sales_to_listings_pct
## 1    83     1533              5.414220
## 2   108     1586              6.809584
## 3   182     1689             10.775607
## 4   200     1708             11.709602
## 5   202     1771             11.405985
## 6   189     1803             10.482529
summary(case$sales_to_listings_pct)
##    Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
##   5.014   8.980  10.963  11.874  13.492  38.713

The median sales-to-active-listings ratio is approximately 10.96%.

The central 50% of observations range from approximately 8.98% to 13.49%, while the maximum value is approximately 38.71%.

Higher values indicate a greater number of monthly sales relative to the number of active listings.

However, this indicator should not be interpreted as a direct conversion rate for individual property listings because the properties sold during a month are not necessarily the same properties included in the active listing count.

For this reason, the ratio can indicate market activity relative to inventory, but it cannot directly establish marketing effectiveness.

Data such as campaign expenditure, property views, enquiries, listing dates and time on market would be required for a more direct evaluation of marketing performance.

Conditional Analysis

Conditional analysis is performed by city, year and calendar month.

Sales by City

tapply(
  case$sales,
  case$city,
  mean
)
##              Beaumont Bryan-College Station                 Tyler 
##              177.3833              205.9667              269.7500 
##         Wichita Falls 
##              116.0667
tapply(
  case$sales,
  case$city,
  sd
)
##              Beaumont Bryan-College Station                 Tyler 
##              41.48395              84.98374              61.96380 
##         Wichita Falls 
##              22.15192

Tyler has the highest mean monthly sales, approximately 269.75, while Wichita Falls has the lowest, approximately 116.07.

Bryan-College Station has the greatest absolute variability in monthly sales, with a standard deviation of approximately 84.98.

Sales by Year

tapply(
  case$sales,
  case$year,
  mean
)
##     2010     2011     2012     2013     2014 
## 168.6667 164.1250 186.1458 211.9167 230.6042
tapply(
  case$sales,
  case$year,
  sd
)
##     2010     2011     2012     2013     2014 
## 60.53708 63.87042 70.90509 83.99641 95.51490

Mean monthly sales per city decrease slightly from approximately 168.67 in 2010 to 164.13 in 2011.

They subsequently increase, reaching approximately 230.60 in 2014.

The annual standard deviations also increase over the period.

These standard deviations combine differences between cities and differences between months, rather than representing variation within a single city alone.

Sales by Month

tapply(
  case$sales,
  case$month,
  mean
)
##      1      2      3      4      5      6      7      8      9     10     11 
## 127.40 140.85 189.45 211.70 238.85 243.55 235.75 231.45 182.35 179.90 156.85 
##     12 
## 169.40
tapply(
  case$sales,
  case$month,
  sd
)
##        1        2        3        4        5        6        7        8 
## 43.38372 51.06783 59.17812 65.40489 83.11582 94.99832 96.27421 79.22883 
##        9       10       11       12 
## 72.51807 74.95395 55.46670 60.74658

Mean sales are lowest in January, at approximately 127.40, and highest in June, at approximately 243.55.

Sales are also relatively high from May to August, suggesting the presence of a seasonal pattern.

July has the greatest dispersion across cities and years, with a standard deviation of approximately 96.27.

mean_sales_by_month <-
  tapply(
    case$sales,
    case$month,
    mean
  )

barplot(
  mean_sales_by_month,
  main = "Mean Sales by Calendar Month (2010-2014)",
  xlab = "Month",
  ylab = "Mean Monthly Sales per City",
  col = "steelblue"
)

Visualisations with ggplot2

Distribution of Median Prices by City

ggplot(
  case,
  aes(
    x = city,
    y = median_price
  )
) +
  geom_boxplot(fill = "steelblue") +
  labs(
    title = "Distribution of Monthly Median Prices by City",
    x = "City",
    y = "Monthly Median Sale Price ($)"
  ) +
  theme_minimal() +
  theme(
    axis.text.x = element_text(
      angle = 25,
      hjust = 1
    )
  )

Bryan-College Station has the highest median of monthly median prices, approximately $155,400, followed by Tyler, Beaumont and Wichita Falls.

The respective values are approximately:

  • Bryan-College Station: $155,400
  • Tyler: $142,200
  • Beaumont: $130,750
  • Wichita Falls: $102,300

The boxplots show clear differences in price levels between cities, although some overlap exists between their distributions.

Some monthly observations are identified as outliers according to the standard boxplot rule. These observations are not necessarily data errors, but rather unusually high or low monthly median prices.

It is also important to note that these boxplots represent distributions of monthly median prices, not the prices of individual properties.

Distribution of Sales Volume by City

ggplot(
  case,
  aes(
    x = city,
    y = volume
  )
) +
  geom_boxplot(fill = "steelblue") +
  labs(
    title = "Distribution of Monthly Sales Volume by City",
    x = "City",
    y = "Monthly Sales Volume (Million $)"
  ) +
  theme_minimal() +
  theme(
    axis.text.x = element_text(
      angle = 25,
      hjust = 1
    )
  )

Tyler has the highest median monthly sales volume, approximately $45.08 million.

Wichita Falls has the lowest median monthly sales volume, approximately $13.71 million, and also shows relatively limited dispersion.

Bryan-College Station shows the greatest dispersion in monthly sales volume, with a standard deviation of approximately $17.25 million and an interquartile range of approximately $23.72 million.

Distribution of Sales Volume by Year

ggplot(
  case,
  aes(
    x = factor(year),
    y = volume
  )
) +
  geom_boxplot(fill = "steelblue") +
  labs(
    title = "Distribution of Monthly Sales Volume by Year",
    x = "Year",
    y = "Monthly Sales Volume (Million $)"
  ) +
  theme_minimal()

Median monthly sales volume increases particularly during 2013 and 2014, reaching approximately $36.83 million in 2014.

Dispersion also increases over the observed period.

Each annual boxplot combines observations from all four cities and therefore reflects both geographical and seasonal differences.

Stacked Sales by Month and City

ggplot(
  case,
  aes(
    x = factor(month),
    y = sales,
    fill = city
  )
) +
  geom_col() +
  labs(
    title = "Total Sales by Calendar Month and City (2010-2014)",
    x = "Month",
    y = "Total Number of Sales",
    fill = "City"
  ) +
  theme_minimal()

geom_col() uses the actual values of the sales variable as the heights of the bars.

Because the bars are stacked, observations corresponding to the same month are added together.

By contrast, geom_bar() would count the number of rows by default rather than use the existing sales values.

Total sales are highest in June, with 4,871 sales, and lowest in January, with 2,548 sales.

Tyler contributes the largest number of sales in every calendar month.

Normalised Stacked Sales

ggplot(
  case,
  aes(
    x = factor(month),
    y = sales,
    fill = city
  )
) +
  geom_col(position = "fill") +
  scale_y_continuous(
    labels = function(x) {
      paste0(round(x * 100), "%")
    }
  ) +
  labs(
    title = "City Shares of Sales by Calendar Month (2010-2014)",
    x = "Month",
    y = "Share of Sales",
    fill = "City"
  ) +
  theme_minimal()

In this graph, every bar is normalised to 100%.

The coloured sections therefore represent the proportion of sales attributable to each city during each calendar month.

Tyler has the highest share in every month.

Bryan-College Station generally has a larger contribution between May and August than during the winter months.

Unlike the previous stacked bar chart, this visualisation compares relative shares rather than absolute sales totals.

Monthly Sales by City and Year

Year is added to the analysis through separate panels using facet_wrap(). This avoids overcrowding the graph while allowing the same visual structure to be compared across different years.

ggplot(
  case,
  aes(
    x = factor(month),
    y = sales,
    fill = city
  )
) +
  geom_col() +
  facet_wrap(
    ~ year,
    ncol = 2
  ) +
  labs(
    title = "Monthly Sales by City and Year",
    x = "Month",
    y = "Total Number of Sales",
    fill = "City"
  ) +
  theme_minimal()

Sales generally increase during the middle months of each year.

Overall sales levels are higher in 2013 and especially in 2014.

Using the same vertical scale across panels makes it possible to compare absolute sales levels between years.

Conclusions and Recommendations

Main Findings

The analysis identifies several relevant patterns in the historical data.

After a slight decline in 2011, overall sales activity increases considerably.

annual_total_sales <-
  tapply(
    case$sales,
    case$year,
    sum
  )

annual_total_sales
##  2010  2011  2012  2013  2014 
##  8096  7878  8935 10172 11069

Total annual sales increase from 7,878 in 2011 to 11,069 in 2014, corresponding to growth of approximately 40.5%.

Tyler records the highest overall sales activity, with an average of approximately 269.75 property sales per month.

It also records the highest median monthly sales volume, approximately $45.08 million.

Bryan-College Station has the highest typical price level, with a median of monthly median prices of approximately $155,400.

It also displays substantial sales growth and relatively high variability in both monthly sales and monthly sales volume.

Beaumont shows an overall increase in sales during the period, despite an initial decline in 2011.

Wichita Falls has the lowest sales activity and price levels and does not show a comparable sustained increase in sales.

A clear seasonal pattern is also visible.

Across cities and years, average sales are lowest in January and highest in June, while market activity remains relatively high between May and August.

Sales volume is the quantitative variable with the greatest relative variability, with a coefficient of variation of approximately 53.71%.

It is also the variable with the strongest asymmetry, showing positive skewness caused by some relatively high-value city-month observations.

The sales-to-active-listings ratio has a median of approximately 10.96% and provides an indicator of sales activity relative to available inventory.

However, it cannot directly measure the effectiveness of individual property listings or marketing campaigns.

Recommendations for Texas Realty Insights

Based on the historical results, Tyler and Bryan-College Station may deserve further investigation as potential priority markets.

Tyler combines high sales activity with high total sales volume, while Bryan-College Station combines relatively high typical prices with substantial sales growth.

However, any current investment or business decision should first be supported by updated market data and should also consider acquisition costs, competition and profitability.

The historical seasonal pattern suggests that listings and sales resources could be prepared in advance of the stronger May-August period.

Before increasing marketing expenditure, however, the company should verify whether the same seasonal pattern remains present in more recent data.

Pricing strategies should also reflect differences between local markets.

The observed differences in monthly median prices suggest that a single pricing approach would not be appropriate for all cities.

Property characteristics and local comparable sales should also be considered when establishing prices for individual listings.

Sales, active listings and months of inventory should be monitored together.

A relatively high sales-to-listings ratio combined with lower inventory duration may indicate a more active market, although this relationship should be investigated statistically rather than assumed.

To evaluate marketing effectiveness more directly, Texas Realty Insights should collect additional information such as:

  • marketing expenditure
  • listing views
  • customer enquiries
  • listing dates
  • time on market
  • final sale outcomes

These variables would allow a more direct evaluation of marketing performance and cost-effectiveness.

Limitations

The dataset contains aggregated monthly observations rather than information on individual properties.

Consequently, differences in prices may reflect changes in the types of properties sold as well as genuine changes in market conditions.

Sales values and prices are expressed in nominal dollars and have not been adjusted for inflation.

The analysis is descriptive.

It identifies historical patterns and statistical relationships but does not establish their causes and cannot guarantee that the same patterns will continue in the future.

Updated data, together with detailed information on individual properties and marketing activities, would strengthen future analyses and support more informed business decisions.