Real data is rarely linearly separable, so SVM uses the kernel trick to bend the boundary: a linear kernel draws a straight frontier, while polynomial and radial (RBF) kernels draw curved, non-linear ones;
Use an SVM when you want a strong binary (or multi-class) classifier and you suspect the boundary between classes may be non-linear;
SVM shines on medium-sized datasets with many predictors, copes well with a non-linear frontier via the kernel, and is robust because it leans only on the support vectors.
“ANTECESSOR”: Maximal Margin Classifier e o Hiperplano (ver o livro);
E X E M P L O
E X E M P L O
O ALUNO ATENTO DEVE ESTAR SE PERGUNTANDO
Por que dosagem²? E não dosagem³? ou
\[
\frac{\pi}{4} \times \sqrt{\text{Dosagem}}
\] AFINAL, COMO DECIDIMOS COMO TRANSFORMAR OS DADOS????
KERNEL FUNCTIONS E KERNEL TRICK
Trabalha com produtos internos ENTRE OS PARES E NÃO COM TRANSFORMAÇÃO DIMENSIONAL;
Por isso ele é otimizado;
SISTEMATICAMENTE acham o Support Vector Classifier mesmo MUITAS DIMENSÕES;
Linear: boa referência inicial; frequente em dados de texto.
Polinomial: incorpora interações entre atributos.
Radial Basis Function: permite relações locais e fronteiras flexíveis.
SUPPORT VECTOR CLASSIFIER
Rather than seeking the largest possible margin so thatevery observation is not only on the correct side of the hyperplane but also on the correct side of the margin, we instead allow some observations to be on the incorrect side of the margin, or even the incorrect side of the hyperplane.
The margin is soft because it can be violated by some of the training observations;
O treinamento considera todos os exemplos. Ao encontrar a solução, alguns ficam com coeficientes nulos e outros passam a integrar diretamente a função de decisão.
EXEMPLO PRÁTICO
You have a clinical dataset and a yes/no question: will a patient test positive for diabetes, given a handful of measurements — glucose, BMI, age, blood pressure, and so on?
The support vector machine (SVM) is built for exactly this. Its idea is geometric and intuitive: among all the boundaries that separate the two classes, pick the one with the widest margin;
the boundary that keeps the biggest possible gap on either side. A wide margin is a confident, robust boundary, and it depends only on the handful of points sitting right on the edge of that gap.
Those border points are the support vectors, and they give the method its name.
BANCO DE DADOS
768 women of Pima Indian heritage, with the goal of predicting diabetes (a binary pos/neg outcome) from eight clinical predictors — number of pregnancies (pregnant), plasma glucose, blood pressure, triceps skinfold thickness, serum insulin, body mass index (BMI), the diabetes pedigree function, and age.
Separação Treino e Teste
library(rsample)data("PimaIndiansDiabetes2", package ="mlbench")# Drop rows with missing measurements (SVM needs complete cases)pima <-na.omit(PimaIndiansDiabetes2)set.seed(123)split <-initial_split(pima, prop =0.80, strata = diabetes)train_data <-training(split)test_data <-testing(split)c(complete =nrow(pima), train =nrow(train_data), test =nrow(test_data))
complete train test
392 313 79
# The stratified split preserves the class balanceprop.table(table(train_data$diabetes))
About 313 patients to learn from, 79 held back to be the judge, and 8 predictors. Roughly a third of the patients are diabetes-positive — a moderately imbalanced but workable classification problem.
O QUE O SVM REALMENTE FAZ? MARGENS E VETORES DE SUPORTE DESTAS MARGENS
O QUE O SVM REALMENTE FAZ? MARGENS E VETORES DE SUPORTE DESTAS MARGENS
Radial-kernel SVM on just two predictors — glucose and mass (BMI)
The boundary only needs the points near it — the support vectors. Points deep inside their own region could be deleted without moving the line at all; only the border cases matter. That is the whole economy of an SVM: a boundary defined by a small, decisive set of points, with the widest possible margin around it.
MESMA COISA SÓ QUE MUDANDO O KERNEL PARA LINEAR
Setting default kernel parameters
TUNING THE COST WITH CROSS-VALIDATION
C: it determines the number and severity of the vio-lations to the margin (and to the hyperplane) that we will tolerate. We canthink of C as a budget for the amount that the margin can be violatedby the n observations.
IMPORTÂNCIA DA REGULARIZAÇÃO E DA NORMALIZAÇÃO: SVM CALCULA DISTÂNCIA E NÃO PROBABILIDADE.
Small C → a soft, wide margin. The model tolerates some misclassified training points in exchange for a simpler boundary that usually generalizes better.
Large C → a hard, narrow margin. The model strains to classify every training point correctly, which can overfit the training noise.
TUNING THE COST WITH CROSS-VALIDATION
We score each candidate C by 10-fold cross-validation on the training set, on accuracy and ROC AUC, then keep the best.
Here the chosen C is small (around 0.125) — the soft, forgiving margin generalizes best on this data. That is a common finding: on noisy real data, a wider margin beats a tighter one.
The tuning curve: accuracy vs cost
The tuning curve: accuracy vs cost
The most useful diagnostic when tuning an SVM is cross-validated accuracy against the cost.
For a small cost (left) the margin is soft and wide, and cross-validated accuracy is at its best.
As the cost grows (right) the margin tightens, the model starts chasing individual training points, and accuracy drifts down and flattens.
The dashed line marks the cross-validation winner — the sweet spot between an over-soft margin (underfit) and an over-hard one (overfit).
Accuracy ≈ 0.76 — the model gets the diabetes call right about 76% of the time on patients it never saw during fitting. Accuracy can flatter a model on an imbalanced outcome.
ROC AUC ≈ 0.85 — the probability that the model scores a random positive patient higher than a random negative one. 0.5 is a coin flip, 1.0 is perfect; 0.85 is a genuinely good classifier.
dat <-data.frame(x = x, y =as.factor(y))library(e1071)
Attaching package: 'e1071'
The following object is masked from 'package:tune':
tune
The following object is masked from 'package:parsnip':
tune
The following object is masked from 'package:rsample':
permutations
svmfit <-svm( y ~ .,data = dat,kernel ="linear",cost =10,scale =FALSE)
O HIPERPARÂMETRO (C) controla o trade-off entre maximizar a margem e penalizar erros de classificação.
The argument scale = FALSE tells the svm() function not to scale each feature to have mean zero or standard deviation one; depending on the application, one might prefer to use scale = TRUE.
SVM()
plot(svmfit , dat)
SVM()
svmfit$index
[1] 1 2 5 7 14 16 17
summary(svmfit)
Call:
svm(formula = y ~ ., data = dat, kernel = "linear", cost = 10, scale = FALSE)
Parameters:
SVM-Type: C-classification
SVM-Kernel: linear
cost: 10
Number of Support Vectors: 7
( 4 3 )
Number of Classes: 2
Levels:
-1 1
HIPERPARAMETRIZANDO
svmfit <-svm(y ~ ., data = dat , kernel ="linear",cost =0.1, scale =FALSE )plot(svmfit , dat)