Support Vector Machines

Máquinas de Vetores de Suporte

M.Sc. Renato Barreira

SUPORT VECTOR MACHINE

  1. Aprendizado de Máquina Supervisionado;
  2. Real data is rarely linearly separable, so SVM uses the kernel trick to bend the boundary: a linear kernel draws a straight frontier, while polynomial and radial (RBF) kernels draw curved, non-linear ones;
  3. Use an SVM when you want a strong binary (or multi-class) classifier and you suspect the boundary between classes may be non-linear;
  4. SVM shines on medium-sized datasets with many predictors, copes well with a non-linear frontier via the kernel, and is robust because it leans only on the support vectors.

“ANTECESSOR”: Maximal Margin Classifier e o Hiperplano (ver o livro);

E X E M P L O

E X E M P L O

O ALUNO ATENTO DEVE ESTAR SE PERGUNTANDO

Por que dosagem²? E não dosagem³? ou

\[ \frac{\pi}{4} \times \sqrt{\text{Dosagem}} \] AFINAL, COMO DECIDIMOS COMO TRANSFORMAR OS DADOS????

KERNEL FUNCTIONS E KERNEL TRICK

  1. Trabalha com produtos internos ENTRE OS PARES E NÃO COM TRANSFORMAÇÃO DIMENSIONAL;
  2. Por isso ele é otimizado;
  3. SISTEMATICAMENTE acham o Support Vector Classifier mesmo MUITAS DIMENSÕES;
  • Linear: boa referência inicial; frequente em dados de texto.
  • Polinomial: incorpora interações entre atributos.
  • Radial Basis Function: permite relações locais e fronteiras flexíveis.

SUPPORT VECTOR CLASSIFIER

  1. Rather than seeking the largest possible margin so thatevery observation is not only on the correct side of the hyperplane but also on the correct side of the margin, we instead allow some observations to be on the incorrect side of the margin, or even the incorrect side of the hyperplane.
  2. The margin is soft because it can be violated by some of the training observations;
  3. O treinamento considera todos os exemplos. Ao encontrar a solução, alguns ficam com coeficientes nulos e outros passam a integrar diretamente a função de decisão.

EXEMPLO PRÁTICO

  1. You have a clinical dataset and a yes/no question: will a patient test positive for diabetes, given a handful of measurements — glucose, BMI, age, blood pressure, and so on?

  2. The support vector machine (SVM) is built for exactly this. Its idea is geometric and intuitive: among all the boundaries that separate the two classes, pick the one with the widest margin;

  3. the boundary that keeps the biggest possible gap on either side. A wide margin is a confident, robust boundary, and it depends only on the handful of points sitting right on the edge of that gap.

  4. Those border points are the support vectors, and they give the method its name.

BANCO DE DADOS

768 women of Pima Indian heritage, with the goal of predicting diabetes (a binary pos/neg outcome) from eight clinical predictors — number of pregnancies (pregnant), plasma glucose, blood pressure, triceps skinfold thickness, serum insulin, body mass index (BMI), the diabetes pedigree function, and age.

Separação Treino e Teste

library(rsample)
data("PimaIndiansDiabetes2", package = "mlbench")

# Drop rows with missing measurements (SVM needs complete cases)
pima <- na.omit(PimaIndiansDiabetes2)

set.seed(123)
split      <- initial_split(pima, prop = 0.80, strata = diabetes)
train_data <- training(split)
test_data  <- testing(split)

c(complete = nrow(pima), train = nrow(train_data), test = nrow(test_data))
complete    train     test 
     392      313       79 
# The stratified split preserves the class balance
prop.table(table(train_data$diabetes))

      neg       pos 
0.6677316 0.3322684 

BANCO DE DADOS

  pregnant glucose pressure triceps insulin mass pedigree age diabetes
4        1      89       66      23      94 28.1    0.167  21      neg
5        0     137       40      35     168 43.1    2.288  33      pos
7        3      78       50      32      88 31.0    0.248  26      pos
9        2     197       70      45     543 30.5    0.158  53      pos

About 313 patients to learn from, 79 held back to be the judge, and 8 predictors. Roughly a third of the patients are diabetes-positive — a moderately imbalanced but workable classification problem.

O QUE O SVM REALMENTE FAZ? MARGENS E VETORES DE SUPORTE DESTAS MARGENS

O QUE O SVM REALMENTE FAZ? MARGENS E VETORES DE SUPORTE DESTAS MARGENS

Radial-kernel SVM on just two predictors — glucose and mass (BMI)

The boundary only needs the points near it — the support vectors. Points deep inside their own region could be deleted without moving the line at all; only the border cases matter. That is the whole economy of an SVM: a boundary defined by a small, decisive set of points, with the widest possible margin around it.

MESMA COISA SÓ QUE MUDANDO O KERNEL PARA LINEAR

 Setting default kernel parameters  

TUNING THE COST WITH CROSS-VALIDATION

C: it determines the number and severity of the vio-lations to the margin (and to the hyperplane) that we will tolerate. We canthink of C as a budget for the amount that the margin can be violatedby the n observations.

IMPORTÂNCIA DA REGULARIZAÇÃO E DA NORMALIZAÇÃO: SVM CALCULA DISTÂNCIA E NÃO PROBABILIDADE.

# A tibble: 1 × 2
   cost .config              
  <dbl> <chr>                
1 0.125 Preprocessor1_Model01

Small C → a soft, wide margin. The model tolerates some misclassified training points in exchange for a simpler boundary that usually generalizes better.

Large C → a hard, narrow margin. The model strains to classify every training point correctly, which can overfit the training noise.

TUNING THE COST WITH CROSS-VALIDATION

We score each candidate C by 10-fold cross-validation on the training set, on accuracy and ROC AUC, then keep the best.

Here the chosen C is small (around 0.125) — the soft, forgiving margin generalizes best on this data. That is a common finding: on noisy real data, a wider margin beats a tighter one.

The tuning curve: accuracy vs cost

The tuning curve: accuracy vs cost

The most useful diagnostic when tuning an SVM is cross-validated accuracy against the cost.

For a small cost (left) the margin is soft and wide, and cross-validated accuracy is at its best.

As the cost grows (right) the margin tightens, the model starts chasing individual training points, and accuracy drifts down and flattens.

The dashed line marks the cross-validation winner — the sweet spot between an over-soft margin (underfit) and an over-hard one (overfit).

Finalizando o Exemplo:

# A tibble: 3 × 4
  .metric     .estimator .estimate .config             
  <chr>       <chr>          <dbl> <chr>               
1 accuracy    binary         0.759 Preprocessor1_Model1
2 roc_auc     binary         0.845 Preprocessor1_Model1
3 brier_class binary         0.158 Preprocessor1_Model1

Accuracy ≈ 0.76 — the model gets the diabetes call right about 76% of the time on patients it never saw during fitting. Accuracy can flatter a model on an imbalanced outcome.

ROC AUC ≈ 0.85 — the probability that the model scores a random positive patient higher than a random negative one. 0.5 is a coin flip, 1.0 is perfect; 0.85 is a genuinely good classifier.

EXEMPLO DO LIVRO (MAIS SIMPLES)

set.seed (1)
x <- matrix(rnorm (20 * 2), ncol = 2)
y <- c(rep(-1, 10), rep(1, 10))
x[y == 1, ] <- x[y == 1, ] + 1
plot(x, col = (3 - y))

USANDO O SVM()

dat <- data.frame(x = x, y = as.factor(y))

library(e1071)

Attaching package: 'e1071'
The following object is masked from 'package:tune':

    tune
The following object is masked from 'package:parsnip':

    tune
The following object is masked from 'package:rsample':

    permutations
svmfit <- svm(
  y ~ .,
  data = dat,
  kernel = "linear",
  cost = 10,
  scale = FALSE
)

O HIPERPARÂMETRO (C) controla o trade-off entre maximizar a margem e penalizar erros de classificação.

The argument scale = FALSE tells the svm() function not to scale each feature to have mean zero or standard deviation one; depending on the application, one might prefer to use scale = TRUE.

SVM()

plot(svmfit , dat)

SVM()

svmfit$index
[1]  1  2  5  7 14 16 17
summary(svmfit)

Call:
svm(formula = y ~ ., data = dat, kernel = "linear", cost = 10, scale = FALSE)


Parameters:
   SVM-Type:  C-classification 
 SVM-Kernel:  linear 
       cost:  10 

Number of Support Vectors:  7

 ( 4 3 )


Number of Classes:  2 

Levels: 
 -1 1

HIPERPARAMETRIZANDO

svmfit <- svm(y ~ ., data = dat , kernel = "linear",cost = 0.1, scale = FALSE )
plot(svmfit , dat)
svmfit$index
 [1]  1  2  3  4  5  7  9 10 12 13 14 15 16 17 18 20

VALIDAÇÃO CRUZADA

options(scipen = 999)

set.seed(1)

tune.out <- e1071::tune(
  e1071::svm,
  y ~ .,
  data = dat,
  kernel = "linear",
  ranges = list(
    cost = c(0.001, 0.01, 0.1, 1, 5, 10, 100)
  )
)

VALIDAÇÃO CRUZADA

summary(tune.out)

Parameter tuning of 'e1071::svm':

- sampling method: 10-fold cross validation 

- best parameters:
 cost
  0.1

- best performance: 0.05 

- Detailed performance results:
     cost error dispersion
1   0.001  0.55  0.4377975
2   0.010  0.55  0.4377975
3   0.100  0.05  0.1581139
4   1.000  0.15  0.2415229
5   5.000  0.15  0.2415229
6  10.000  0.15  0.2415229
7 100.000  0.15  0.2415229
bestmod <- tune.out$best.model

PREDICT()

xtest <- matrix(rnorm (20 * 2), ncol = 2)
ytest <- sample(c(-1, 1), 20, rep = TRUE)
xtest[ytest == 1, ] <- xtest[ytest == 1, ] + 1
testdat <- data.frame(x = xtest , y = as.factor(ytest))
ypred <- predict(bestmod , testdat)
table( predict = ypred , truth = testdat $y)
       truth
predict -1 1
     -1  9 1
     1   2 8
svmfit <- svm(y ~ ., data = dat , kernel = "linear",cost = .01, scale = FALSE)
ypred <- predict(svmfit , testdat)
table( predict = ypred , truth = testdat $y)
       truth
predict -1  1
     -1 11  6
     1   0  3

KERNEL POYNOMIAL AGORA

set.seed (1)
x <- matrix(rnorm (200 * 2), ncol = 2)
x[1:100 , ] <- x[1:100 , ] + 2
x[101:150 , ] <- x[101:150 , ] - 2
y <- c(rep(1, 150) , rep(2, 50))
dat <- data.frame(x = x, y = as.factor(y))
plot(x, col = y)

TREINO E TESTE

train <- sample (200, 100)
svmfit <- svm(y ~ ., data = dat[train , ], kernel = "radial",gamma = 1, cost = 1)
plot(svmfit , dat[train , ])

MAIS RIGIDEZ NO CUSTO

svmfit <- svm(y ~ ., data = dat[train , ], kernel = "radial",gamma = 1, cost = 1e5)
plot(svmfit , dat[train , ])

MAS NÃO SOMOS ARBITRÁRIOS VAMOS USAR CROSS VALIDATION

set.seed (1)
tune.out <- e1071::tune(svm , y ~ ., data = dat[train , ],kernel = "radial",ranges = list(cost = c(0.1, 1, 10, 100, 1000) ,gamma = c(0.5, 1, 2, 3, 4)))

HIPERPARÂMETRO GAMMA:

Gamma pequeno: Cada observação influencia região maior (fronteira mais suava e simples e risco de UNDERFITTING)

Gamma grande: Cada observação influencia região menor (fronteira mais irregular e complexa e maior risco de OVERFITTING)

RESULTADO CROSS VALIDATION

summary(tune.out)

Parameter tuning of 'svm':

- sampling method: 10-fold cross validation 

- best parameters:
 cost gamma
    1   0.5

- best performance: 0.07 

- Detailed performance results:
     cost gamma error dispersion
1     0.1   0.5  0.26 0.15776213
2     1.0   0.5  0.07 0.08232726
3    10.0   0.5  0.07 0.08232726
4   100.0   0.5  0.14 0.15055453
5  1000.0   0.5  0.11 0.07378648
6     0.1   1.0  0.22 0.16193277
7     1.0   1.0  0.07 0.08232726
8    10.0   1.0  0.09 0.07378648
9   100.0   1.0  0.12 0.12292726
10 1000.0   1.0  0.11 0.11005049
11    0.1   2.0  0.27 0.15670212
12    1.0   2.0  0.07 0.08232726
13   10.0   2.0  0.11 0.07378648
14  100.0   2.0  0.12 0.13165612
15 1000.0   2.0  0.16 0.13498971
16    0.1   3.0  0.27 0.15670212
17    1.0   3.0  0.07 0.08232726
18   10.0   3.0  0.08 0.07888106
19  100.0   3.0  0.13 0.14181365
20 1000.0   3.0  0.15 0.13540064
21    0.1   4.0  0.27 0.15670212
22    1.0   4.0  0.07 0.08232726
23   10.0   4.0  0.09 0.07378648
24  100.0   4.0  0.13 0.14181365
25 1000.0   4.0  0.15 0.13540064

TESTE PREDICT:

table(true = dat[-train , "y"],pred = predict(tune.out$best.model , newdata = dat[-train , ]))
    pred
true  1  2
   1 67 10
   2  2 21

MAS E SE EU TIVER MUITAS CLASSES????

set.seed (1)
x <- rbind(x, matrix(rnorm (50 * 2), ncol = 2))
y <- c(y, rep(0, 50))
x[y == 0, 2] <- x[y == 0, 2] + 2
dat <- data.frame(x = x, y = as.factor(y))
par(mfrow = c(1, 1))> plot(x, col = (y + 1))
logical(0)

RESULTADO MULTICLASSES

svmfit <- svm(y ~ ., data = dat , kernel = "radial",cost = 10, gamma = 1)
plot(svmfit , dat)

CONCLUINDO:

  1. MODELO DE ML muito engenhoso que serve para tudo, incluindo imagem e texto;
  2. Bancos de Dados menores;
  3. Multidimensionalidade “infinita”;
  4. NECESSIDADE DE NORMALIZAR, REGULARIZAR E ESCALAR OS DADOS;
  5. BOA EXPLICABILIDADE.

FALTOU:

A MATEMÁTICA POR TRÁS DOS KERNELS Ela é relativamente fácil, vale a pena dar uma olhada depois.

FONTES:

Conceitos e exemplos: https://www.datanovia.com/learn/machine-learning/classification/support-vector-machine

Conceitos e exemplos: An Introduction to Statistical Learning with Applications in R.

Conceitos e exemplos: Statquest with Josh Starmer https://www.youtube.com/watch?v=efR1C6CvhmE