2024-06-10

Classification

  • Classification is a type of supervised machine learning algorithm which predicts the “class” of a given input.

  • One of the most common types of classification algorithms is binary classification in which an input is either classified as either a “0” or “1”.

  • For example, we could create a binary classification algorithm to classify an input as either a clementine, “0”, or an orange, “1”.

  • Logistic regression is generally a great model for classification and is implemented using the logistic or sigmoid function

    • \(g(z) = \frac{1}{1+e^{-z}}\)
      • where \(0 < g(z) < 1\)

Sigmoid Function

  • \(g(z) = \frac{1}{1+e^{-z}}\)
  • Note, the z values used to create the plot were in the domain [-10,10] yet all of the sigmoid function’s outputs are between 0 & 1

Logistic Regression Model

  • A machine learning model, \(f_{(\vec{w}, \ b)}(\vec{x})\), is implemented using weights * biases
    • \(\vec{x}\) = input vector holding data to classify
    • \(\vec{w}\) = weight vector
      • used to attribute a ‘weight’ or level of significance to certain features
      • Ex. features for the Clementine vs Orange classifier can be the fruit’s weight, color, size, etc.
    • \(b\) = bias
  • A logistic regression classification model is defined as:
    • \(f_{(\vec{w}, \ b)}(\vec{x}) = g(\vec{w} \cdot \vec{x} + b) = \frac{1}{1 + e^{-(\vec{w} \cdot \vec{x} + b)}}\)
    • note, \(z = \vec{w} \cdot \vec{x} + b\)

Threshold & Decision Boundary

  • Since the sigmoid function maps inputs to a value between 0 and 1, a threshold value must be established to classify outputs as either a 0 or a 1
    • Ex. If we set threshold = 0.5 & \(f(x) \geq 0.5\)
      • then the model prediction is \(\hat{y} = 1\) or that the input \(x\) is an orange
  • The predicted outputs can be plotted on a feature space and the separation between those belonging to different classes can be shown by a decision boundary
  • Decision boundaries are often used to visualize model performance on various test sets

Linear Decision Boundary

Cost Functions & Gradient Descent

  • Cost functions, \(j(w,b)\), are used to evaluate how good the current values for the model’s parameters, \(w\) & \(b\), are at predicting the target value \(y\)
    • since the cost function is proportional to the avg. error, our goal is to find values for \(w\) & \(b\) that minimize the cost function
  • Gradient descent is an optimization algorithm used to minimize the cost function
    • works by identifying the shortest path towards \(j\)’s closest local minima corresponding to the starting point
    • update weights: \(w_j = w_j - \alpha[\frac{1}{m}\sum_{i=1}^{m}(\hat{y^{(i)}}-y^{(i)})x_j^{(i)}]\)
    • update bias: \(b = b - \alpha[\frac{1}{m}\sum_{i=1}^{m}(\hat{y^{(i)}}-y^{(i)})]\)

Loss Function

  • A common cost function for gradient descent is the loss function
  • \(j_{(\vec{w}, b)} = \frac{1}{m}\sum_{i=1}^{m}L(f_{\vec{w},b}(\vec{x}^{(i)}), \ y^{(i)})\)
    • if \(y^{(i)}=1\)
      • \(L = - log(f_{\vec{w},b}(\vec{x}^{(i)}))\)
    • if \(y^{(i)}=0\)
      • \(L = - log(1 - f_{\vec{w},b}(\vec{x}^{(i)}))\)

Running Gradient Descent

  • As the number of gradient descent iteration increases, the loss or error between the prediction, \(\hat{y}\), and target \(y\) decreases
  • Minimal loss means better model performance!