Support Vector Machines (SVM): Definition and Examples
By Dr. Zubair Khalid, DVM, MS, PhD ·

Support vector machines (SVM) are supervised learning models that separate classes by finding the decision boundary with the widest possible gap between them. The training points that sit closest to that boundary are the support vectors, and they alone define it. This article explains the margin, walks through a small worked example, and shows when SVM classification is a good choice.
Quick Answer
- An SVM draws a boundary (a line in 2D, a hyperplane in higher dimensions) that separates classes with the largest possible margin.
- The margin is the distance between the boundary and the nearest points of each class. Wider margins usually generalize better.
- Only the closest points, called support vectors, determine the boundary. Other points can move without changing it [1].
- Kernels let an SVM draw curved boundaries by implicitly mapping data into a higher-dimensional space [1].
- SVMs work well when you have many features relative to samples, but training cost grows quickly with the number of training points [1].
What Support Vector Machines Mean
In plain terms, an SVM is a classifier that asks: what is the widest street I can draw between these two groups of points? The center line of that street is the decision boundary. New points are classified by which side of the line they fall on.
The precise definition: a support vector machine is a supervised learning algorithm that finds the maximum-margin hyperplane separating labeled classes, where the hyperplane is determined by a subset of training points (the support vectors) and can be made nonlinear through a kernel function [1].
Two ideas carry most of the weight. First, maximizing the margin. Among all boundaries that separate the classes perfectly, the SVM picks the one with the most clearance on both sides. Second, sparsity. Because only the support vectors matter, the fitted model is memory efficient and can be stored as a small subset of the training data [1].
How It Works
For linearly separable data, the decision function is:
$$f(x) = w \cdot x + b$$
Each symbol means:
- $x$ is the input vector (the features of one observation).
- $w$ is the weight vector. It is perpendicular to the decision boundary and sets its orientation.
- $b$ is the bias (intercept). It shifts the boundary away from the origin.
- $f(x) = 0$ defines the boundary itself. Positive values fall in one class, negative values in the other.
The margin width is:
$$\text{margin} = \frac{2}{\lVert w \rVert}$$
where $\lVert w \rVert$ is the length of the weight vector. To widen the margin, the algorithm shrinks $\lVert w \rVert$. Training solves a quadratic programming problem that balances two goals: keep points on the correct side, and keep $\lVert w \rVert$ small [1].
The points that sit exactly on the margin edge are the support vectors. They satisfy $|f(x)| = 1$, and they are the only points that enter the final formula for $w$ [1].
Real data is rarely perfectly separable, so a regularization parameter $C$ controls the trade-off between a wide margin and allowing some points to violate it. A large $C$ penalizes violations heavily and fits the training data tightly. A small $C$ accepts more violations in exchange for a wider margin [2].
When a straight boundary cannot separate the classes, a kernel function replaces the dot product $w \cdot x$ with a similarity measure computed in a higher-dimensional space. Common kernels include linear, polynomial, and radial basis function (RBF) kernels, and you can also supply a custom kernel [1].
Worked Example
The dataset is 12 synthetic points in two dimensions, split into two linearly separable classes of six points each.
| feature1 | feature2 | class |
|---|---|---|
| 1.0 | 1.5 | -1 |
| 1.5 | 1.0 | -1 |
| 2.0 | 2.5 | -1 |
| 2.5 | 2.0 | -1 |
| 3.0 | 3.5 | -1 |
| 3.5 | 3.0 | -1 |
| 5.0 | 5.5 | 1 |
| 5.5 | 5.0 | 1 |
| 6.0 | 6.5 | 1 |
| 6.5 | 6.0 | 1 |
| 7.0 | 7.5 | 1 |
| 7.5 | 7.0 | 1 |
The steps and their computed values:
- Dataset size: n = 12 points, 2 features, 2 classes.
- Weight vector: w = [0.5, 0.5].
- Bias: b = -4.25.
- Decision function: f(x) = 0.5·x1 + 0.5·x2 + (-4.25), so the boundary is the line x1 + x2 = 8.5.
- Margin width: 2 / ||w|| = 2 / 0.7071 = 2.8284.
- Support vectors found: 2 points at indices [5, 7], the points (3.5, 3.0) and (5.5, 5.0), each with |f(x)| = 1.
The code below fits a linear SVM with a very large value of $C$, which forces a hard margin on separable data.
import numpy as np
from sklearn.svm import SVC
X = np.array([[1.0,1.5],[1.5,1.0],[2.0,2.5],[2.5,2.0],
[3.0,3.5],[3.5,3.0],[5.0,5.5],[5.5,5.0],
[6.0,6.5],[6.5,6.0],[7.0,7.5],[7.5,7.0]])
y = np.array([-1,-1,-1,-1,-1,-1,1,1,1,1,1,1])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
print(clf.coef_, clf.intercept_, clf.support_) # inspect weights, bias, support vector indices
Output:
[[0.5 0.5]] [-4.25] [5 7]
The fitted boundary has slope -1 and sits exactly halfway between the two clusters, on the line x1 + x2 = 8.5. The margin of 2.83 units is the total clearance between the two nearest points on either side. The solver reports two support vectors, (3.5, 3.0) from class -1 and (5.5, 5.0) from class 1, and these alone define the boundary. The points (3.0, 3.5) and (5.0, 5.5) also lie exactly on the margin edges, but the solver needs only one point on each side here.
How to Interpret It
Read the weights as direction, not as importance scores. In this example, $w_1 = 0.5$ and $w_2 = 0.5$ are equal, so the boundary runs at 45 degrees and both features contribute equally to the separation. If one weight were much larger, the boundary would tilt toward being perpendicular to that feature.
Read the margin as a confidence buffer. A margin of 2.83 units means a new point must be more than about 1.41 units from the boundary before it is clearly on one side. Points that land inside the margin are ambiguous, and points on the wrong side are misclassified.
Read the sign of $f(x)$ as the predicted class. Plug a point into the decision function. A positive result means class 1, a negative result means class -1. The magnitude tells you how far the point sits from the boundary, though it is not a probability. SVMs do not produce probability estimates directly, and scikit-learn computes them through an expensive five-fold cross-validation when you request them [1].
To judge overall performance, pair the SVM with a threshold-free metric. The ROC curve and AUC give a clearer picture of ranking quality than raw accuracy on an imbalanced set.
When to Use It (and when not to)
Use an SVM when:
- The number of features is large relative to the number of samples. SVMs remain effective in high-dimensional settings [1].
- You need a nonlinear boundary and want to choose the kernel rather than engineer features by hand [1].
- Memory matters. The decision function uses only the support vectors, so prediction is cheap [1].
- You want a method with solid theory behind it. For background, see Introduction to Statistical Learning.
Avoid or reconsider an SVM when:
- Your training set is very large. Compute and storage requirements rise rapidly with the number of training vectors [1].
- You need calibrated probabilities out of the box. SVMs require extra cross-validation work for that [1].
- You have many more features than samples and are unsure about kernel choice. Overfitting risk rises, and regularization and kernel selection need care [1].
If your data is low-dimensional and you want a simple, interpretable baseline, K-Nearest Neighbors is often easier to explain. If you want class probabilities directly, see Bayesian classifiers and Naive Bayes.
SVM vs Logistic Regression
Both draw a linear boundary, but they optimize different things. Logistic regression maximizes the likelihood of the labels. An SVM maximizes the margin and ignores points far from the boundary.
| Aspect | SVM | Logistic Regression |
|---|---|---|
| Objective | Maximize margin between classes | Maximize likelihood of observed labels |
| Influence of far points | None, only support vectors matter [1] | Every point contributes to the fit |
| Nonlinear boundaries | Yes, via kernels [1] | Requires manual feature expansion |
| Probability output | Not direct, needs extra cross-validation [1] | Yes, by design |
| High-dimensional data | Strong performer [1] | Can work, but needs regularization |
Common Mistakes
- Forgetting to scale features. SVMs measure distance, so a feature on a 0 to 1000 scale will dominate one on a 0 to 1 scale. Standardize or normalize your features first.
- Tuning C without a validation set. A very large C fits training data tightly and can overfit. Use cross-validation to pick C instead of guessing.
- Assuming every training point shapes the boundary. Only support vectors do [1]. If you want to understand a fitted model, inspect the support vectors, not the whole dataset.
- Reading decision function values as probabilities. A value of 2.0 is not "twice as confident" as 1.0. Request probabilities explicitly if you need them, and expect the extra cost [1].
- Using a nonlinear kernel on data that is already linearly separable. It adds complexity with no benefit. Start with a linear kernel and move up only if the fit is poor.
- Ignoring the training size. An SVM on hundreds of thousands of rows can be slow to train. Check whether a simpler model gets you close enough first [1].
Limitations
SVMs do not tell you why a point was classified a certain way. The kernel trick makes the boundary expressive but opaque, and the weights in a nonlinear model do not map cleanly onto individual features. If you need explanations per prediction, this is a real obstacle.
Training cost is the other hard limit. The underlying quadratic programming solver scales between $O(n_{features} \times n_{samples}^2)$ and higher, so doubling your training rows can multiply training time several times over [1]. SVMs also assume your classes can be separated reasonably well in the chosen feature space. If the kernel is wrong, no amount of tuning C will fix the boundary. For a broader view of where this fits, see Statistical Analysis Methods.
Frequently Asked Questions
What is a support vector in simple terms?
A support vector is a training point that sits closest to the decision boundary. These points define where the boundary goes. Remove a point far from the boundary and nothing changes. Remove a support vector and the boundary shifts [1].
Can SVM handle more than two classes?
Yes. The standard approach is to break a multi-class problem into several two-class problems and combine the results. scikit-learn handles this internally, so you pass all your class labels and the library manages the rest [1].
Does SVM work with text data?
Yes, and it is a common choice for text classification because text produces many features relative to the number of documents, which suits SVMs [1]. Scale or normalize your features first, since raw word counts vary widely in magnitude.
What does the C parameter actually do?
C controls how much you penalize points that fall on the wrong side of the margin. A large C means few violations are tolerated and the margin shrinks. A small C allows more violations and produces a wider margin [2]. Tune it with cross-validation.
Is SVM the same as support vector regression?
No. They share the same underlying machinery, but classification predicts a class label while support vector regression predicts a continuous value. The regression variant uses an epsilon tube instead of a margin between classes [2]. If you want a regression baseline instead, see What is Regression Analysis?.
References
Further Reading
- Estimating Average Treatment Effects with Support Vector Machines
- Lever J, Krzywinski M, Altman N (2016). Classification evaluation. Nature Methods
- scikit-learn User Guide
- Wilson G, Bryan J, Cranston K et al. (2017). Good enough practices in scientific computing. PLOS Computational Biology
Related Articles
- Python for Machine Learning: A Beginner's Guide
- Multivariate Analysis: Definition, Methods and Examples
- K-Nearest Neighbors (KNN): Algorithm and Examples
- Bivariate Data: Definition, Examples and Analysis
- K-Means Clustering: How It Works With a Worked Example
- Introduction To Statistical Learning
- Elements Of Statistical Learning
- Principal Component Analysis (PCA) in Biological Research