GATE DA Machine Learning Formula Sheet with Worked Lines
The machine learning formulas you need for GATE DA cover regression, classification, decision trees, neural networks, model selection, clustering and PCA. This sheet gives each with its condition, a one-line example of my own and the trap beside it. Every formula is checked against the last-minute sheet in the GATE DA 2027 book.
In this guide
Key takeaways
- Least squares, ridge and PCA are linear algebra: learn the matrix forms, not only the scalar ones.
- Ridge shrinks coefficients towards zero but never exactly to zero.
- Logistic regression, LDA with a shared covariance and the perceptron all draw linear boundaries.
- Entropy in bits uses ; the logistic loss uses .
- Small in k-nearest neighbour means high variance; large means high bias.
- Count network parameters layer by layer, and check whether biases are included.
The terms this sheet uses
The training data are pairs . Each is a vector of features, and is its label: a number in regression, a class in classification. is the matrix of features, one row per example.
A model has a weight vector and a bias . The loss measures how wrong the predictions are, and training minimises it. is the learning rate. is the natural logarithm, and is used wherever the answer is in bits.
Regression
| Formula | Watch out for |
|---|---|
| ; ; | The line passes through |
| Through the origin: | No centring here |
| Simple regression: | |
| , so | Fitted values , is a projection |
| Ridge: minimise ; | Invertible for every |
| Larger : smaller coefficients, more bias, less variance | Never exactly zero; is least squares |
Examples: for and , and . Through the origin with and , ; ridge with gives . If and , then .
Trap: The intercept is usually not penalised in ridge regression. Shrinking it would move the fitted line away from the data's mean.
Classification: logistic, Bayes, LDA and SVM
| Formula | Watch out for |
|---|---|
| ; | The boundary is linear |
| Loss ; gradient | No closed-form solution |
| Naive Bayes: | Features independent given the class |
| binary features, two classes: parameters | likelihoods plus one prior, since the two priors sum to 1 |
| Laplace: | Removes zero probabilities |
| Bayes error at : | Choose the largest posterior |
| LDA with a shared covariance: linear boundary | Unequal covariances give a quadratic one |
| Fisher: maximise ; two classes | At most directions for classes |
| Perceptron: if , , | Converges iff linearly separable |
| SVM margin ; , only on support vectors | Large allows little slack |
Examples: if , then . Ten binary features with two classes need 21 naive Bayes parameters. A value never seen in 8 examples of a class with 2 possible values gets . With , the SVM margin is .
The k-nearest neighbour rule takes a majority vote of the closest points. With the training error is zero. Scale the features first, and pick an odd to avoid ties between two classes. XOR is not linearly separable, so no perceptron learns it; one hidden layer does.
Trees and evaluation
| Formula | Watch out for |
|---|---|
| 1 bit at 50-50, 0 when pure | |
| 0.5 at 50-50 for two classes | |
| Gain | Weight children by size |
| Accuracy ; precision ; recall | Do not swap FP and FN |
| ; specificity | is a harmonic mean |
Examples: a parent with 4 positive and 4 negative examples has . Split into and , each child has , so the gain is about 0.189. With , and : precision 0.8, recall , .
The book's last-minute sheet covers all seven technical sections in this form, each result with its condition beside it. It is in the GATE DA 2027 book, with 907 questions with worked solutions and 10 full mock tests.
Neural networks and model selection
| Formula | Watch out for |
|---|---|
| ; ReLU | Derivative 1 for , 0 for |
| Sigmoid | At most |
| No non-linearity: is one linear layer | Depth alone adds nothing |
| Weights ; biases | No bias on the input layer |
| Errors flow back through times | |
| ; with : | Without the , a factor 2 appears |
| Expected squared error | Flexible models: low bias, high variance |
| k-fold: models, each on folds; leave-one-out: models on points | Never tune on the test set |
Examples: a 4–3–1 network has weights and 19 parameters. One step with , , , and loss gives and . Five-fold validation on 100 points trains 5 models, each on 80 points.
Remember: Regularisation trades a little bias for less variance. That one sentence answers most bias–variance statements about ridge, in k-NN and tree depth.
Clustering and PCA
| Formula | Watch out for |
|---|---|
| k-means minimises | Never increases; local optimum depends on the start |
| k-medoids: centres are data points, any dissimilarity | Less sensitive to outliers |
| Agglomerative: merges from singletons | Divisive works top-down |
| Single linkage: minimum pair distance; complete: maximum; average: mean | Single linkage chains |
| Euclidean ; Manhattan | The metric can change the first merge |
| PCA: eigenvectors of on centred ; share | Centre first; scaling matters |
| Scores ; via SVD, | Components are orthogonal |
Examples: from to the Euclidean distance is 5 and the Manhattan distance 7. Eigenvalues 6, 3 and 1 give shares of 0.6, 0.3 and 0.1.
In one line: PCA is an eigenvalue problem on the centred covariance matrix, so the linear algebra sheet does half the work.
Using the sheet in the exam
Entropy, and sigmoid values end in decimals. Practise them on the GATE virtual calculator. When a set of statements comes as a multiple-select question, judge each one separately. The MCQ, MSQ and NAT marking scheme, common to every GATE paper, gives an MSQ no partial credit and no negative marks. If you know the CS syllabus, GATE CS vs GATE DA shows that machine learning is DA's own territory. How a raw mark becomes a score is in GATE score calculation and normalisation.
More formula sheets: all of GATE DA · DBMS and Algorithms · Linear Algebra and Calculus · Probability and Statistics
Quick revision
- , and the line passes through .
- Ridge: , shrinking but never to zero.
- Naive Bayes multiplies the prior by each feature's likelihood given the class.
- SVM margin is ; only support vectors matter.
- Entropy uses ; Gini is .
- Precision divides by , recall by .
- Weights , plus one bias per non-input unit.
- PCA share of component is .
Frequently asked questions
What is the formula for information gain in a decision tree?
Information gain is the parent's entropy minus the weighted entropy of the children: , with . Entropy is 1 bit for a 50-50 split of two classes and 0 for a pure node. Use , because the answer is expected in bits.
How many parameters does a fully connected neural network have?
For layers of sizes , count weights, then add one bias for every unit after the input layer. A 4-3-1 network has weights and 4 biases, so 19 parameters. Questions often ask for weights only, so read whether biases are included.
What is the margin of a hard-margin SVM?
The hard-margin SVM minimises subject to , and the width of the margin is . The support vectors lie exactly on . Removing a point that is not a support vector leaves the solution unchanged, a fact worth knowing for statement-based questions.
What is the difference between k-fold and leave-one-out cross-validation?
In k-fold cross-validation you split the data into folds, train models each on folds, and average the validation scores. Leave-one-out is the case : it trains models, each on points. It is nearly unbiased but expensive and has high variance. Never tune on the test set.
How do you compute the variance explained by a principal component?
Centre the data and find the eigenvalues of its covariance matrix. The variance along component is its eigenvalue , and the share explained is . With eigenvalues 6, 3 and 1, the first component explains 60 per cent and the first two explain 90 per cent. PCA is sensitive to feature scaling.
Sources
Dates, fees and the syllabus are set by the GATE 2027 organising institute and can change. Always confirm at gate2027.iitm.ac.in.