Estimating a Probability

Time of arrival of individuals A and B Individual A arrives at a random time between 10 and 11:30pm and B arrives at a random time between 10:30 and 12:00am

What is the probability that individual B arrives before A?

What is the expected difference in the two arrival times?

We assume A and B are independent and uniformly distributed on times (10.5 and 12) for B and (10 to 11.5) for A

The probability is expressed as P(B < A)

We simulate a large number of values (1000) from the distribution of (A,B) independently

A = runif(1000, 10, 11.5)

B = runif(1000, 10.5, 12)

The probability P(B < A) is estimated by the proportion of simulated pairs (A,B) where A is a smaller than B. We count the number of pairs where A < B by the sum function and divide this by the total number of simulations

prob = sum(B < A)/1000
prob
## [1] 0.213

The estimated probability that B arrives before A P(B < A) is 0.225

plot(A, B)
polygon(c(10.5, 11.5, 11.5, 10.5),
        c(10.5, 10.5, 11.5, 10.5), density = 10, angle = 135)

The Monte Carlo estimate of the probability is the proportion of points that fall in the shaded region.

Let’s now compute the standard error (SE) of a proportion of this estimate

sqrt(prob * (1 - prob)/1000)
## [1] 0.01294724

Applying a normal approximation for the sampling distribution of p there is a 95% confidence that the probability B will arrive earlier than A is between

0.225 - 1.96 * 0.0132
## [1] 0.199128
0.225 + 1.96 * 0.0132
## [1] 0.250872

What is the expected difference in the two arrival times?

E(B - A)

difference = B - A

The Monte Carlo estimate of E(B - A) is the mean of these differences. The estimated standard error of this sample mean estimate is the standard deviation of the differences divided by the square root of the simulation sample size

mc.est = mean(difference)
se.est = sd(difference)/sqrt(1000)
c(mc.est, se.est)
## [1] 0.52169499 0.01964135

So we estimate that B will arrive 0.5 hours later than A; since the standard error is only 0.02 hours, there is a 95% confidence that the true difference is within 0.04 hours of this estimate

Introduction to Markov Chains

As described by Wikipedia:

A Markov chain is a stochastic model describing a sequence of possible events in which the probability of each event depends only on the state attained in the previous event.

This property of memorylessness (dependence only on the preceding state) is called the Markov property.

A Markov chain is used to model a random process consisting of multiple states. These systems can be finite or infinite. An example for a finite Markov Process could be the position of an elevator in a building. The number of states (floors) is finite and we are interested in the probability to find the system in a specific state at a future time. Infinite Markov Processes have (as the name suggest) an infinite number of states. We have knowledge about the transition between the states and are interested in the state of the system at a future time.

Tracking the movement of particles to another location will depend only on their current location. We’ve determined the following probabilities for the movement of a particle:

Of the particles 30% will remain stationary, 30% will move to the Middle, while the remaining 40% will go to the End. Of all the particles in the Middle, 50% and 30% will move to up and down, respectively; 20% will remain in Middle. Particles in the Middle have a 40% chance of moving to up, 40% chance of staying down; 20% will move to Middle. Once a particle is in an area, it can either move to the next area or stay back in the same area. This movement will be decided only by the current location and not the sequence of past locations. The location in this example includes up, down and Middle. It follows all the properties of Markov Chains because the current location has the power to predict the next location.

location <- c("Up","Middle","Down")
locationTPM <- matrix(c(0.3,0.3,0.4,
                        0.5,0.2,0.3,
                        0.4,0.2,0.4), 
                        nrow=3, byrow=T, dimnames = list(location,location))
locationTPM
##         Up Middle Down
## Up     0.3    0.3  0.4
## Middle 0.5    0.2  0.3
## Down   0.4    0.2  0.4
#install.packages("markovchain")
library(markovchain)
## Loading required package: Matrix
## Package:  markovchain
## Version:  0.10.3
## Date:     2026-02-02 06:30:37 UTC
## BugReport: https://github.com/spedygiorgio/markovchain/issues
newlocation <- new("markovchain",
                   states = location,
                   transitionMatrix = locationTPM,
                   name = "Particle Movement")
newlocation
## Particle Movement 
##  A  3 - dimensional discrete Markov Chain defined by the following states: 
##  Up, Middle, Down 
##  The transition matrix  (by rows)  is defined as follows: 
##         Up Middle Down
## Up     0.3    0.3  0.4
## Middle 0.5    0.2  0.3
## Down   0.4    0.2  0.4

Question: Assuming that a particle is currently in the up location, what is the probability that the particle will again be in the up location after two moves?

A particle can reach the up again in the second movement in three different ways:

Staying in the same location Probability: P(N-N) = 0.30.3 = 0.09

Middle to up Probability: P(M-N) = 0.40.4 = 0.12

down to up Probability: P(S-N) = 0.5*0.4 = 0.2

Therefore, Probability of second movement tup = 0.09 + 0.12 + 0.20 = 0.41 (0.40 approximately)

newlocation^2
## Particle Movement^2 
##  A  3 - dimensional discrete Markov Chain defined by the following states: 
##  Up, Middle, Down 
##  The transition matrix  (by rows)  is defined as follows: 
##          Up Middle Down
## Up     0.40   0.23 0.37
## Middle 0.37   0.25 0.38
## Down   0.38   0.24 0.38

This gives us the direct probability of a particle coming back up after two movements.

The calculation can be done for subsequent movements as well. The probability of coming back up in 4th movement

newlocation^3
## Particle Movement^3 
##  A  3 - dimensional discrete Markov Chain defined by the following states: 
##  Up, Middle, Down 
##  The transition matrix  (by rows)  is defined as follows: 
##           Up Middle  Down
## Up     0.383  0.240 0.377
## Middle 0.388  0.237 0.375
## Down   0.386  0.238 0.376

However if we increase n (No. of movements) the predictive power tends to go down, where the Markov Chain reaches an equilibrium called stationary state. In the above case, after 9 movements the state becomes stationary -

newlocation^9
## Particle Movement^9 
##  A  3 - dimensional discrete Markov Chain defined by the following states: 
##  Up, Middle, Down 
##  The transition matrix  (by rows)  is defined as follows: 
##               Up    Middle      Down
## Up     0.3853211 0.2385321 0.3761468
## Middle 0.3853211 0.2385321 0.3761468
## Down   0.3853211 0.2385321 0.3761468
steadyStates(newlocation)
##             Up    Middle      Down
## [1,] 0.3853211 0.2385321 0.3761468

So we can say that, in the stationary state, a particle has a probability of 0.385 of ending up.

If there were 100 movements in all and each completes 9 moves in a day, how many of them will end in the Middle?

steadyStates(newlocation)*100
##            Up   Middle     Down
## [1,] 38.53211 23.85321 37.61468

So, around 38 particles will end up, 24 in the Middle and 37 down.