Texas Realty Insights
This project analyses historical real estate market data for four
Texas cities between 2010 and 2014.
The objective is to identify sales trends, evaluate variability and
distributional characteristics, investigate seasonal patterns, compare
cities and periods, and provide statistical insights that may support
business decisions.
case <- read.csv("realestate_texas.csv")
library(ggplot2)
str(case)
## 'data.frame': 240 obs. of 8 variables:
## $ city : chr "Beaumont" "Beaumont" "Beaumont" "Beaumont" ...
## $ year : int 2010 2010 2010 2010 2010 2010 2010 2010 2010 2010 ...
## $ month : int 1 2 3 4 5 6 7 8 9 10 ...
## $ sales : int 83 108 182 200 202 189 164 174 124 150 ...
## $ volume : num 14.2 17.7 28.7 26.8 28.8 ...
## $ median_price : num 163800 138200 122400 123200 123100 ...
## $ listings : int 1533 1586 1689 1708 1771 1803 1857 1830 1829 1779 ...
## $ months_inventory: num 9.5 10 10.6 10.6 10.9 11.1 11.7 11.6 11.7 11.5 ...
Variable Analysis
The dataset contains 240 observations covering four cities over 60
months, from January 2010 to December 2014.
Each row represents one city in one specific month and year. The
dataset therefore combines a geographical dimension with a time
dimension.
The variables can be classified as follows:
- city: nominal categorical variable. Frequency
distributions and comparisons between cities are appropriate.
- year: discrete time variable. It can be used to
compare annual results and identify trends.
- month: discrete time variable with a cyclical
structure. It can be used to investigate seasonal patterns. Month
numbers should not be interpreted as ordinary quantitative
measurements.
- sales: discrete quantitative variable representing
the number of properties sold in each city-month.
- volume: continuous quantitative variable
representing the total value of sales in millions of US dollars.
- median_price: continuous quantitative variable
representing the median property sale price for each city-month.
- listings: discrete quantitative variable
representing the number of active property listings.
- months_inventory: continuous quantitative variable
indicating the estimated number of months required to sell the current
inventory.
Position, variability and shape measures are appropriate for the
quantitative variables sales, volume,
median_price, listings and
months_inventory.
To correctly analyse the chronological evolution of the market, year
and month are combined into a single date variable.
case$date <- as.Date(
sprintf("%04d-%02d-01", case$year, case$month)
)
colSums(is.na(case))
## city year month sales
## 0 0 0 0
## volume median_price listings months_inventory
## 0 0 0 0
## date
## 0
sum(duplicated(case[c("city", "year", "month")]))
## [1] 0
There are no missing values and no duplicated combinations of city,
year and month.
Measures of Position, Variability and Shape
Frequency distributions are appropriate for the categorical and time
variables.
table(case$city)
##
## Beaumont Bryan-College Station Tyler
## 60 60 60
## Wichita Falls
## 60
table(case$year)
##
## 2010 2011 2012 2013 2014
## 48 48 48 48 48
table(case$month)
##
## 1 2 3 4 5 6 7 8 9 10 11 12
## 20 20 20 20 20 20 20 20 20 20 20 20
Each city has 60 observations, each year contains 48 observations and
each calendar month contains 20 observations.
This confirms that all four cities are represented throughout the
same five-year period.
Sales
summary(case$sales)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 79.0 127.0 175.5 192.3 247.0 423.0
sd(case$sales)
## [1] 79.65111
var(case$sales)
## [1] 6344.3
IQR(case$sales)
## [1] 120
sd(case$sales) / mean(case$sales) * 100
## [1] 41.42203
Monthly sales per city have a mean of approximately
192.29 and a median of 175.50.
The standard deviation is approximately 79.65, while
the interquartile range is 120.
The mean is higher than the median, which is consistent with a
positively skewed distribution.
Other Quantitative Variables
summary(
case[c(
"volume",
"median_price",
"listings",
"months_inventory"
)]
)
## volume median_price listings months_inventory
## Min. : 8.166 Min. : 73800 Min. : 743 Min. : 3.400
## 1st Qu.:17.660 1st Qu.:117300 1st Qu.:1026 1st Qu.: 7.800
## Median :27.062 Median :134500 Median :1618 Median : 8.950
## Mean :31.005 Mean :132665 Mean :1738 Mean : 9.193
## 3rd Qu.:40.893 3rd Qu.:150050 3rd Qu.:2056 3rd Qu.:10.950
## Max. :83.547 Max. :180000 Max. :3296 Max. :14.900
sapply(
case[c(
"volume",
"median_price",
"listings",
"months_inventory"
)],
sd
)
## volume median_price listings months_inventory
## 16.651447 22662.148687 752.707756 2.303669
sapply(
case[c(
"volume",
"median_price",
"listings",
"months_inventory"
)],
var
)
## volume median_price listings months_inventory
## 2.772707e+02 5.135730e+08 5.665690e+05 5.306889e+00
sapply(
case[c(
"volume",
"median_price",
"listings",
"months_inventory"
)],
IQR
)
## volume median_price listings months_inventory
## 23.2335 32750.0000 1029.5000 3.1500
Monthly sales volume averages approximately $31.01
million, with a median of approximately $27.06
million and a standard deviation of approximately
$16.65 million.
The mean of the monthly median prices is approximately
$132,665.42, while their median is
$134,500. The standard deviation is approximately
$22,662.15.
Active listings average approximately 1,738.02, with
a median of 1,618.50 and a standard deviation of
approximately 752.71.
Months of inventory average approximately 9.19
months, with a median of 8.95 months and a
standard deviation of approximately 2.30 months.
These statistics combine differences between cities, seasonal
variation and changes over time. For this reason, they should also be
interpreted together with the conditional analyses presented later.
Coefficients of Variation
The coefficient of variation allows variables with different units
and scales to be compared in relative terms.
cv_values <- sapply(
case[c(
"sales",
"volume",
"median_price",
"listings",
"months_inventory"
)],
function(x) sd(x) / mean(x) * 100
)
round(cv_values, 2)
## sales volume median_price listings
## 41.42 53.71 17.08 43.31
## months_inventory
## 25.06
Skewness
The following calculation measures the asymmetry of each quantitative
distribution.
skewness_values <- sapply(
case[c(
"sales",
"volume",
"median_price",
"listings",
"months_inventory"
)],
function(x) mean(((x - mean(x)) / sd(x))^3)
)
round(skewness_values, 3)
## sales volume median_price listings
## 0.714 0.879 -0.362 0.645
## months_inventory
## 0.041
Sales show positive skewness, meaning that relatively high monthly
sales observations extend the upper tail of the distribution.
Sales volume also has positive skewness and therefore presents some
relatively high-value city-month observations.
Median price has modest negative skewness.
Listings have positive skewness, indicating the presence of some
relatively high listing counts.
Months of inventory have skewness close to zero, indicating
relatively little overall asymmetry.
Variables with the Greatest Variability and Asymmetry
round(cv_values, 2)
## sales volume median_price listings
## 41.42 53.71 17.08 43.31
## months_inventory
## 25.06
names(cv_values)[which.max(cv_values)]
## [1] "volume"
round(skewness_values, 3)
## sales volume median_price listings
## 0.714 0.879 -0.362 0.645
## months_inventory
## 0.041
names(skewness_values)[which.max(abs(skewness_values))]
## [1] "volume"
The variable with the greatest relative variability is
volume, with a coefficient of variation of
approximately 53.71%.
The coefficient of variation is more appropriate than directly
comparing variances or standard deviations because the variables have
different units and scales.
Sales volume is also the variable with the greatest absolute
skewness, approximately 0.879.
Its positive skewness indicates that some city-month observations
have particularly high total sales values.
By comparison, months_inventory has skewness close to
zero, while median_price has moderate negative
skewness.
Sales Classes, Frequencies and Gini Heterogeneity
The quantitative variable sales is divided into five
classes.
case$sales_class <- cut(
case$sales,
breaks = c(0, 100, 200, 300, 400, 500),
labels = c(
"0-99",
"100-199",
"200-299",
"300-399",
"400-499"
),
right = FALSE
)
class_freq <- table(case$sales_class)
class_freq
##
## 0-99 100-199 200-299 300-399 400-499
## 20 127 67 23 3
class_prop <- class_freq / sum(class_freq)
round(class_prop, 3)
##
## 0-99 100-199 200-299 300-399 400-499
## 0.083 0.529 0.279 0.096 0.013
The class frequencies are:
- 0-99: 20 observations
- 100-199: 127 observations
- 200-299: 67 observations
- 300-399: 23 observations
- 400-499: 3 observations
The 100-199 class is the most frequent, containing
approximately 52.9% of all observations.
barplot(
class_freq,
main = "Frequency Distribution of Sales Classes",
xlab = "Number of Sales",
ylab = "Number of Observations",
col = "steelblue"
)

The Gini heterogeneity index is calculated as follows:
gini_heterogeneity <- 1 - sum(class_prop^2)
round(gini_heterogeneity, 3)
## [1] 0.626
The Gini heterogeneity index is approximately
0.626.
A value of zero would indicate that all observations belong to a
single class.
With five classes, the maximum possible value is
0.800, which would occur if all classes contained equal
proportions of observations.
The observed value therefore indicates that the observations are
distributed across several classes, although more than half are
concentrated in the 100-199 sales class.
This measure is an index of heterogeneity between categories and
should not be confused with the Gini coefficient commonly used to
measure inequality.
Its value also depends on the selected class boundaries.
Probability Analysis
The probability calculations are based on the relative frequency of
rows in the dataset.
prob_beaumont <- sum(case$city == "Beaumont") / nrow(case)
prob_july <- sum(case$month == 7) / nrow(case)
prob_december_2012 <- sum(
case$year == 2012 & case$month == 12
) / nrow(case)
prob_beaumont
## [1] 0.25
prob_july
## [1] 0.08333333
prob_december_2012
## [1] 0.01666667
The probability that a randomly selected row refers to
Beaumont is:
60 / 240 = 0.25 = 25%
The probability that a randomly selected row refers to
July is:
20 / 240 = 0.0833 = 8.33%
The probability that a randomly selected row refers to
December 2012 is:
4 / 240 = 0.0167 = 1.67%
These probabilities refer to the frequency of observations in the
dataset and not to the probability of an individual property being
sold.
Creation of New Variables
Average Property Price
The dataset contains total sales volume and the number of sales.
Since volume is expressed in millions of dollars, it is
multiplied by 1,000,000 before being divided by the number of properties
sold.
case$average_price <- case$volume * 1000000 / case$sales
head(
case[c(
"volume",
"sales",
"average_price"
)]
)
## volume sales average_price
## 1 14.162 83 170626.5
## 2 17.690 108 163796.3
## 3 28.701 182 157697.8
## 4 26.819 200 134095.0
## 5 28.833 202 142737.6
## 6 27.219 189 144015.9
summary(case$average_price)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 97010 132939 156588 154320 173915 213234
Average property prices across city-month observations range from
approximately $97,010 to $213,234.
The median monthly average property price is approximately
$156,588, while the unweighted mean across city-month
observations is approximately $154,320.
average_price represents the arithmetic mean sale price
within each city-month.
This differs from median_price, because the arithmetic
mean is more strongly influenced by relatively expensive properties.
The mean of average_price across all rows gives equal
weight to each city-month and is therefore not the same as the
sales-weighted average price of all individual properties sold.
Sales-to-Listings Ratio
A second variable is created to measure monthly sales relative to the
number of active listings.
case$sales_to_listings_pct <-
case$sales / case$listings * 100
head(
case[c(
"sales",
"listings",
"sales_to_listings_pct"
)]
)
## sales listings sales_to_listings_pct
## 1 83 1533 5.414220
## 2 108 1586 6.809584
## 3 182 1689 10.775607
## 4 200 1708 11.709602
## 5 202 1771 11.405985
## 6 189 1803 10.482529
summary(case$sales_to_listings_pct)
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## 5.014 8.980 10.963 11.874 13.492 38.713
The median sales-to-active-listings ratio is approximately
10.96%.
The central 50% of observations range from approximately
8.98% to 13.49%, while the maximum value is
approximately 38.71%.
Higher values indicate a greater number of monthly sales relative to
the number of active listings.
However, this indicator should not be interpreted as a direct
conversion rate for individual property listings because the properties
sold during a month are not necessarily the same properties included in
the active listing count.
For this reason, the ratio can indicate market activity relative to
inventory, but it cannot directly establish marketing effectiveness.
Data such as campaign expenditure, property views, enquiries, listing
dates and time on market would be required for a more direct evaluation
of marketing performance.
Conditional Analysis
Conditional analysis is performed by city, year and calendar
month.
Sales by City
tapply(
case$sales,
case$city,
mean
)
## Beaumont Bryan-College Station Tyler
## 177.3833 205.9667 269.7500
## Wichita Falls
## 116.0667
tapply(
case$sales,
case$city,
sd
)
## Beaumont Bryan-College Station Tyler
## 41.48395 84.98374 61.96380
## Wichita Falls
## 22.15192
Tyler has the highest mean monthly sales,
approximately 269.75, while Wichita
Falls has the lowest, approximately
116.07.
Bryan-College Station has the greatest absolute
variability in monthly sales, with a standard deviation of approximately
84.98.
Sales by Year
tapply(
case$sales,
case$year,
mean
)
## 2010 2011 2012 2013 2014
## 168.6667 164.1250 186.1458 211.9167 230.6042
tapply(
case$sales,
case$year,
sd
)
## 2010 2011 2012 2013 2014
## 60.53708 63.87042 70.90509 83.99641 95.51490
Mean monthly sales per city decrease slightly from approximately
168.67 in 2010 to 164.13 in 2011.
They subsequently increase, reaching approximately 230.60 in
2014.
The annual standard deviations also increase over the period.
These standard deviations combine differences between cities and
differences between months, rather than representing variation within a
single city alone.
Sales by Month
tapply(
case$sales,
case$month,
mean
)
## 1 2 3 4 5 6 7 8 9 10 11
## 127.40 140.85 189.45 211.70 238.85 243.55 235.75 231.45 182.35 179.90 156.85
## 12
## 169.40
tapply(
case$sales,
case$month,
sd
)
## 1 2 3 4 5 6 7 8
## 43.38372 51.06783 59.17812 65.40489 83.11582 94.99832 96.27421 79.22883
## 9 10 11 12
## 72.51807 74.95395 55.46670 60.74658
Mean sales are lowest in January, at approximately
127.40, and highest in June, at
approximately 243.55.
Sales are also relatively high from May to August,
suggesting the presence of a seasonal pattern.
July has the greatest dispersion across cities and years, with a
standard deviation of approximately 96.27.
mean_sales_by_month <-
tapply(
case$sales,
case$month,
mean
)
barplot(
mean_sales_by_month,
main = "Mean Sales by Calendar Month (2010-2014)",
xlab = "Month",
ylab = "Mean Monthly Sales per City",
col = "steelblue"
)

Visualisations with ggplot2
Distribution of Sales Volume by City
ggplot(
case,
aes(
x = city,
y = volume
)
) +
geom_boxplot(fill = "steelblue") +
labs(
title = "Distribution of Monthly Sales Volume by City",
x = "City",
y = "Monthly Sales Volume (Million $)"
) +
theme_minimal() +
theme(
axis.text.x = element_text(
angle = 25,
hjust = 1
)
)

Tyler has the highest median monthly sales volume, approximately
$45.08 million.
Wichita Falls has the lowest median monthly sales volume,
approximately $13.71 million, and also shows relatively
limited dispersion.
Bryan-College Station shows the greatest dispersion in monthly sales
volume, with a standard deviation of approximately $17.25
million and an interquartile range of approximately
$23.72 million.
Distribution of Sales Volume by Year
ggplot(
case,
aes(
x = factor(year),
y = volume
)
) +
geom_boxplot(fill = "steelblue") +
labs(
title = "Distribution of Monthly Sales Volume by Year",
x = "Year",
y = "Monthly Sales Volume (Million $)"
) +
theme_minimal()

Median monthly sales volume increases particularly during
2013 and 2014, reaching approximately $36.83
million in 2014.
Dispersion also increases over the observed period.
Each annual boxplot combines observations from all four cities and
therefore reflects both geographical and seasonal differences.
Stacked Sales by Month and City
ggplot(
case,
aes(
x = factor(month),
y = sales,
fill = city
)
) +
geom_col() +
labs(
title = "Total Sales by Calendar Month and City (2010-2014)",
x = "Month",
y = "Total Number of Sales",
fill = "City"
) +
theme_minimal()

geom_col() uses the actual values of the
sales variable as the heights of the bars.
Because the bars are stacked, observations corresponding to the same
month are added together.
By contrast, geom_bar() would count the number of rows
by default rather than use the existing sales values.
Total sales are highest in June, with 4,871
sales, and lowest in January, with
2,548 sales.
Tyler contributes the largest number of sales in every calendar
month.
Normalised Stacked Sales
ggplot(
case,
aes(
x = factor(month),
y = sales,
fill = city
)
) +
geom_col(position = "fill") +
scale_y_continuous(
labels = function(x) {
paste0(round(x * 100), "%")
}
) +
labs(
title = "City Shares of Sales by Calendar Month (2010-2014)",
x = "Month",
y = "Share of Sales",
fill = "City"
) +
theme_minimal()

In this graph, every bar is normalised to 100%.
The coloured sections therefore represent the proportion of sales
attributable to each city during each calendar month.
Tyler has the highest share in every month.
Bryan-College Station generally has a larger contribution between May
and August than during the winter months.
Unlike the previous stacked bar chart, this visualisation compares
relative shares rather than absolute sales totals.
Monthly Sales by City and Year
Year is added to the analysis through separate panels using
facet_wrap(). This avoids overcrowding the graph while
allowing the same visual structure to be compared across different
years.
ggplot(
case,
aes(
x = factor(month),
y = sales,
fill = city
)
) +
geom_col() +
facet_wrap(
~ year,
ncol = 2
) +
labs(
title = "Monthly Sales by City and Year",
x = "Month",
y = "Total Number of Sales",
fill = "City"
) +
theme_minimal()

Sales generally increase during the middle months of each year.
Overall sales levels are higher in 2013 and
especially in 2014.
Using the same vertical scale across panels makes it possible to
compare absolute sales levels between years.
Historical Sales Trends
To analyse changes over time, the previously created
date variable is used.
ggplot(
case,
aes(
x = date,
y = sales
)
) +
geom_line(
color = "steelblue",
linewidth = 0.7
) +
facet_wrap(
~ city,
ncol = 2
) +
labs(
title = "Monthly Sales Trends by City (2010-2014)",
x = "Year",
y = "Number of Sales"
) +
theme_minimal()

The use of the date variable ensures that the observations are
represented in chronological order.
Separate panels prevent observations from different cities from being
connected by the same line.
Tyler and Bryan-College Station show substantial growth between 2010
and 2014, together with recurring seasonal peaks.
Mean monthly sales in Tyler increase from approximately
227.50 in 2010 to 331.50 in 2014.
For Bryan-College Station, mean monthly sales increase from
approximately 167.58 in 2010 to 260.25 in
2014.
Beaumont also shows overall growth, although sales initially decrease
in 2011.
Its mean monthly sales increase from approximately 156.17 in
2010 to 213.67 in 2014.
Wichita Falls remains at a lower sales level and does not show the
same sustained upward trend.
Its mean monthly sales change from approximately 123.42 in
2010 to 117.00 in 2014.
These results refer specifically to the four cities included in the
dataset during the 2010-2014 period and should not automatically be
generalised to the entire Texas real estate market or to later
periods.
Conclusions and Recommendations
Main Findings
The analysis identifies several relevant patterns in the historical
data.
After a slight decline in 2011, overall sales activity increases
considerably.
annual_total_sales <-
tapply(
case$sales,
case$year,
sum
)
annual_total_sales
## 2010 2011 2012 2013 2014
## 8096 7878 8935 10172 11069
Total annual sales increase from 7,878 in 2011 to
11,069 in 2014, corresponding to growth of
approximately 40.5%.
Tyler records the highest overall sales activity, with an average of
approximately 269.75 property sales per month.
It also records the highest median monthly sales volume,
approximately $45.08 million.
Bryan-College Station has the highest typical price level, with a
median of monthly median prices of approximately
$155,400.
It also displays substantial sales growth and relatively high
variability in both monthly sales and monthly sales volume.
Beaumont shows an overall increase in sales during the period,
despite an initial decline in 2011.
Wichita Falls has the lowest sales activity and price levels and does
not show a comparable sustained increase in sales.
A clear seasonal pattern is also visible.
Across cities and years, average sales are lowest in January and
highest in June, while market activity remains relatively high between
May and August.
Sales volume is the quantitative variable with the greatest relative
variability, with a coefficient of variation of approximately
53.71%.
It is also the variable with the strongest asymmetry, showing
positive skewness caused by some relatively high-value city-month
observations.
The sales-to-active-listings ratio has a median of approximately
10.96% and provides an indicator of sales activity
relative to available inventory.
However, it cannot directly measure the effectiveness of individual
property listings or marketing campaigns.
Recommendations for Texas Realty Insights
Based on the historical results, Tyler and Bryan-College Station may
deserve further investigation as potential priority markets.
Tyler combines high sales activity with high total sales volume,
while Bryan-College Station combines relatively high typical prices with
substantial sales growth.
However, any current investment or business decision should first be
supported by updated market data and should also consider acquisition
costs, competition and profitability.
The historical seasonal pattern suggests that listings and sales
resources could be prepared in advance of the stronger
May-August period.
Before increasing marketing expenditure, however, the company should
verify whether the same seasonal pattern remains present in more recent
data.
Pricing strategies should also reflect differences between local
markets.
The observed differences in monthly median prices suggest that a
single pricing approach would not be appropriate for all cities.
Property characteristics and local comparable sales should also be
considered when establishing prices for individual listings.
Sales, active listings and months of inventory should be monitored
together.
A relatively high sales-to-listings ratio combined with lower
inventory duration may indicate a more active market, although this
relationship should be investigated statistically rather than
assumed.
To evaluate marketing effectiveness more directly, Texas Realty
Insights should collect additional information such as:
- marketing expenditure
- listing views
- customer enquiries
- listing dates
- time on market
- final sale outcomes
These variables would allow a more direct evaluation of marketing
performance and cost-effectiveness.
Limitations
The dataset contains aggregated monthly observations rather than
information on individual properties.
Consequently, differences in prices may reflect changes in the types
of properties sold as well as genuine changes in market conditions.
Sales values and prices are expressed in nominal dollars and have not
been adjusted for inflation.
The analysis is descriptive.
It identifies historical patterns and statistical relationships but
does not establish their causes and cannot guarantee that the same
patterns will continue in the future.
Updated data, together with detailed information on individual
properties and marketing activities, would strengthen future analyses
and support more informed business decisions.