GATE GUIDE

GATE DA Machine Learning Formula Sheet with Worked Lines

By MD ANISH AHAMADUpdated 4 Oct 20266 min read
GATE DA Machine Learning Formula Sheet with Worked Lines

The machine learning formulas you need for GATE DA cover regression, classification, decision trees, neural networks, model selection, clustering and PCA. This sheet gives each with its condition, a one-line example of my own and the trap beside it. Every formula is checked against the last-minute sheet in the GATE DA 2027 book.

In this guide
  1. Key takeaways
  2. The terms this sheet uses
  3. Regression
  4. Classification: logistic, Bayes, LDA and SVM
  5. Trees and evaluation
  6. Neural networks and model selection
  7. Clustering and PCA
  8. Using the sheet in the exam
  9. Quick revision

Key takeaways

The terms this sheet uses

The training data are nn pairs (xi,yi)(x_i, y_i). Each xix_i is a vector of features, and yiy_i is its label: a number in regression, a class in classification. XX is the n×dn \times d matrix of features, one row per example.

A model has a weight vector ww and a bias bb. The loss measures how wrong the predictions are, and training minimises it. η\eta is the learning rate. ln⁡\ln is the natural logarithm, and log⁡2\log_2 is used wherever the answer is in bits.

Regression

Formula Watch out for
y^=β0+β1x\hat{y} = \beta_0 + \beta_1 x; β1=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2\displaystyle \beta_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2}; β0=yˉ−β1xˉ\beta_0 = \bar{y} - \beta_1 \bar{x} The line passes through (xˉ,yˉ)(\bar{x}, \bar{y})
Through the origin: w=∑xiyi∑xi2\displaystyle w = \frac{\sum x_i y_i}{\sum x_i^2} No centring here
R2=1−SSresSStot\displaystyle R^2 = 1 - \frac{SS_{\text{res}}}{SS_{\text{tot}}} Simple regression: R2=r2R^2 = r^2
X⊤Xβ=X⊤yX^\top X\beta = X^\top y, so β^=(X⊤X)−1X⊤y\hat{\beta} = (X^\top X)^{-1}X^\top y Fitted values HyHy, H=X(X⊤X)−1X⊤H = X(X^\top X)^{-1}X^\top is a projection
Ridge: minimise ∥y−Xβ∥2+λ∥β∥2\lVert y - X\beta \rVert^2 + \lambda \lVert \beta \rVert^2; β^=(X⊤X+λI)−1X⊤y\hat{\beta} = (X^\top X + \lambda I)^{-1}X^\top y Invertible for every λ>0\lambda > 0
Larger λ\lambda: smaller coefficients, more bias, less variance Never exactly zero; λ=0\lambda = 0 is least squares

Examples: for x=(1,2,3)x = (1, 2, 3) and y=(2,3,5)y = (2, 3, 5), β1=32=1.5\displaystyle \beta_1 = \frac{3}{2} = 1.5 and β0=103−3=13\displaystyle \beta_0 = \frac{10}{3} - 3 = \frac{1}{3}. Through the origin with x=(1,2)x = (1, 2) and y=(2,3)y = (2, 3), w=85=1.6\displaystyle w = \frac{8}{5} = 1.6; ridge with λ=1\lambda = 1 gives 85+1≈1.33\displaystyle \frac{8}{5 + 1} \approx 1.33. If SSres=20SS_{\text{res}} = 20 and SStot=80SS_{\text{tot}} = 80, then R2=0.75R^2 = 0.75.

Trap: The intercept is usually not penalised in ridge regression. Shrinking it would move the fitted line away from the data's mean.

Classification: logistic, Bayes, LDA and SVM

Formula Watch out for
P(y=1∣x)=σ(w⊤x+b)P(y = 1 \mid x) = \sigma(w^\top x + b); ln⁡p1−p=w⊤x+b\displaystyle \ln\frac{p}{1 - p} = w^\top x + b The boundary is linear
Loss −∑[yiln⁡pi+(1−yi)ln⁡(1−pi)]-\sum \big[y_i \ln p_i + (1 - y_i)\ln(1 - p_i)\big]; gradient ∑(pi−yi)xi\sum (p_i - y_i)x_i No closed-form solution
Naive Bayes: P(C∣x)∝P(C)∏jP(xj∣C)P(C \mid x) \propto P(C)\prod_j P(x_j \mid C) Features independent given the class
KK binary features, two classes: 2K+12K + 1 parameters 2K2K likelihoods plus one prior, since the two priors sum to 1
Laplace: P(xj=v∣C)=count+1NC+number of values of xj\displaystyle P(x_j = v \mid C) = \frac{\text{count} + 1}{N_C + \text{number of values of } x_j} Removes zero probabilities
Bayes error at xx: 1−max⁡CP(C∣x)1 - \max_C P(C \mid x) Choose the largest posterior
LDA with a shared covariance: linear boundary Unequal covariances give a quadratic one
Fisher: maximise J(w)=w⊤SBww⊤SWw\displaystyle J(w) = \frac{w^\top S_B w}{w^\top S_W w}; two classes w∝SW−1(μ1−μ2)w \propto S_W^{-1}(\mu_1 - \mu_2) At most C−1C - 1 directions for CC classes
Perceptron: if yi(w⊤xi+b)≤0y_i(w^\top x_i + b) \le 0, w←w+ηyixiw \leftarrow w + \eta y_i x_i, b←b+ηyib \leftarrow b + \eta y_i Converges iff linearly separable
SVM margin 2∥w∥\displaystyle \frac{2}{\lVert w \rVert}; w=∑αiyixiw = \sum \alpha_i y_i x_i, αi>0\alpha_i > 0 only on support vectors Large CC allows little slack

Examples: if w⊤x+b=2w^\top x + b = 2, then P(y=1∣x)=11+e−2≈0.881\displaystyle P(y = 1 \mid x) = \frac{1}{1 + e^{-2}} \approx 0.881. Ten binary features with two classes need 21 naive Bayes parameters. A value never seen in 8 examples of a class with 2 possible values gets 0+18+2=0.1\displaystyle \frac{0 + 1}{8 + 2} = 0.1. With w=(3,4)w = (3, 4), the SVM margin is 25=0.4\displaystyle \frac{2}{5} = 0.4.

The k-nearest neighbour rule takes a majority vote of the kk closest points. With k=1k = 1 the training error is zero. Scale the features first, and pick an odd kk to avoid ties between two classes. XOR is not linearly separable, so no perceptron learns it; one hidden layer does.

Trees and evaluation

Formula Watch out for
H=−∑kpklog⁡2pkH = -\sum_k p_k \log_2 p_k 1 bit at 50-50, 0 when pure
Gini=1−∑kpk2\text{Gini} = 1 - \sum_k p_k^2 0.5 at 50-50 for two classes
Gain =H(parent)−∑∣child∣∣parent∣H(child)\displaystyle = H(\text{parent}) - \sum \frac{\lvert \text{child} \rvert}{\lvert \text{parent} \rvert}H(\text{child}) Weight children by size
Accuracy TP+TNtotal\displaystyle \frac{TP + TN}{\text{total}}; precision TPTP+FP\displaystyle \frac{TP}{TP + FP}; recall TPTP+FN\displaystyle \frac{TP}{TP + FN} Do not swap FP and FN
F1=2PRP+R\displaystyle F_1 = \frac{2PR}{P + R}; specificity TNTN+FP\displaystyle \frac{TN}{TN + FP} F1F_1 is a harmonic mean

Examples: a parent with 4 positive and 4 negative examples has H=1H = 1. Split into (3+,1−)(3+, 1-) and (1+,3−)(1+, 3-), each child has H≈0.811H \approx 0.811, so the gain is about 0.189. With TP=40TP = 40, FP=10FP = 10 and FN=20FN = 20: precision 0.8, recall 23\displaystyle \frac{2}{3}, F1≈0.727F_1 \approx 0.727.

The book's last-minute sheet covers all seven technical sections in this form, each result with its condition beside it. It is in the GATE DA 2027 book, with 907 questions with worked solutions and 10 full mock tests.

Neural networks and model selection

Formula Watch out for
a=g(w⊤x+b)a = g(w^\top x + b); ReLU g(z)=max⁡(0,z)g(z) = \max(0, z) Derivative 1 for z>0z > 0, 0 for z<0z < 0
Sigmoid σ′=σ(1−σ)\sigma' = \sigma(1 - \sigma) At most 14\displaystyle \frac{1}{4}
No non-linearity: W2W1xW_2 W_1 x is one linear layer Depth alone adds nothing
Weights ∑lnl−1nl\sum_l n_{l-1}n_l; biases ∑l≥1nl\sum_{l \ge 1} n_l No bias on the input layer
∂L∂w=(error at output unit)×(activation at input unit)\displaystyle \frac{\partial L}{\partial w} = (\text{error at output unit}) \times (\text{activation at input unit}) Errors flow back through W⊤W^\top times g′g'
w←w−η∇L(w)w \leftarrow w - \eta \nabla L(w); with 12(y^−y)2\displaystyle \frac{1}{2}(\hat{y} - y)^2: w←w−η(y^−y)xw \leftarrow w - \eta(\hat{y} - y)x Without the 12\displaystyle \frac{1}{2}, a factor 2 appears
Expected squared error =bias2+variance+noise= \text{bias}^2 + \text{variance} + \text{noise} Flexible models: low bias, high variance
k-fold: kk models, each on k−1k - 1 folds; leave-one-out: nn models on n−1n - 1 points Never tune on the test set

Examples: a 4–3–1 network has 12+3=1512 + 3 = 15 weights and 19 parameters. One step with w=1w = 1, x=2x = 2, y=3y = 3, η=0.1\eta = 0.1 and loss 12(y^−y)2\displaystyle \frac{1}{2}(\hat{y} - y)^2 gives y^=2\hat{y} = 2 and w=1−0.1×(−1)×2=1.2w = 1 - 0.1 \times (-1) \times 2 = 1.2. Five-fold validation on 100 points trains 5 models, each on 80 points.

Remember: Regularisation trades a little bias for less variance. That one sentence answers most bias–variance statements about ridge, kk in k-NN and tree depth.

Clustering and PCA

Formula Watch out for
k-means minimises ∑∥x−μc(x)∥2\sum \lVert x - \mu_{c(x)} \rVert^2 Never increases; local optimum depends on the start
k-medoids: centres are data points, any dissimilarity Less sensitive to outliers
Agglomerative: n−1n - 1 merges from nn singletons Divisive works top-down
Single linkage: minimum pair distance; complete: maximum; average: mean Single linkage chains
Euclidean ∑(xi−yi)2\sqrt{\sum (x_i - y_i)^2}; Manhattan ∑∣xi−yi∣\sum \lvert x_i - y_i \rvert The metric can change the first merge
PCA: eigenvectors of S=1nX⊤X\displaystyle S = \frac{1}{n}X^\top X on centred XX; share λk∑jλj\displaystyle \frac{\lambda_k}{\sum_j \lambda_j} Centre first; scaling matters
Scores z=W⊤(x−xˉ)z = W^\top(x - \bar{x}); via SVD, λ=σ2n\displaystyle \lambda = \frac{\sigma^2}{n} Components are orthogonal

Examples: from (0,0)(0, 0) to (3,4)(3, 4) the Euclidean distance is 5 and the Manhattan distance 7. Eigenvalues 6, 3 and 1 give shares of 0.6, 0.3 and 0.1.

In one line: PCA is an eigenvalue problem on the centred covariance matrix, so the linear algebra sheet does half the work.

Using the sheet in the exam

Entropy, F1F_1 and sigmoid values end in decimals. Practise them on the GATE virtual calculator. When a set of statements comes as a multiple-select question, judge each one separately. The MCQ, MSQ and NAT marking scheme, common to every GATE paper, gives an MSQ no partial credit and no negative marks. If you know the CS syllabus, GATE CS vs GATE DA shows that machine learning is DA's own territory. How a raw mark becomes a score is in GATE score calculation and normalisation.

More formula sheets: all of GATE DA · DBMS and Algorithms · Linear Algebra and Calculus · Probability and Statistics

Quick revision

  1. β1=SxySxx\displaystyle \beta_1 = \frac{S_{xy}}{S_{xx}}, and the line passes through (xˉ,yˉ)(\bar{x}, \bar{y}).
  2. Ridge: (X⊤X+λI)−1X⊤y(X^\top X + \lambda I)^{-1}X^\top y, shrinking but never to zero.
  3. Naive Bayes multiplies the prior by each feature's likelihood given the class.
  4. SVM margin is 2∥w∥\displaystyle \frac{2}{\lVert w \rVert}; only support vectors matter.
  5. Entropy uses log⁡2\log_2; Gini is 1−∑pk21 - \sum p_k^2.
  6. Precision divides by TP+FPTP + FP, recall by TP+FNTP + FN.
  7. Weights ∑nl−1nl\sum n_{l-1}n_l, plus one bias per non-input unit.
  8. PCA share of component kk is λk∑jλj\displaystyle \frac{\lambda_k}{\sum_j \lambda_j}.

Frequently asked questions

What is the formula for information gain in a decision tree?

Information gain is the parent's entropy minus the weighted entropy of the children: H(parent)−∑∣child∣∣parent∣H(child)\displaystyle H(\text{parent}) - \sum \frac{\lvert \text{child} \rvert}{\lvert \text{parent} \rvert} H(\text{child}), with H=−∑pklog⁡2pkH = -\sum p_k \log_2 p_k. Entropy is 1 bit for a 50-50 split of two classes and 0 for a pure node. Use log⁡2\log_2, because the answer is expected in bits.

How many parameters does a fully connected neural network have?

For layers of sizes n0,n1,…,nLn_0, n_1, \ldots, n_L, count ∑lnl−1nl\sum_l n_{l-1} n_l weights, then add one bias for every unit after the input layer. A 4-3-1 network has 12+3=1512 + 3 = 15 weights and 4 biases, so 19 parameters. Questions often ask for weights only, so read whether biases are included.

What is the margin of a hard-margin SVM?

The hard-margin SVM minimises 12∥w∥2\displaystyle \frac{1}{2}\lVert w \rVert^2 subject to yi(w⊤xi+b)≥1y_i(w^\top x_i + b) \ge 1, and the width of the margin is 2∥w∥\displaystyle \frac{2}{\lVert w \rVert}. The support vectors lie exactly on yi(w⊤xi+b)=1y_i(w^\top x_i + b) = 1. Removing a point that is not a support vector leaves the solution unchanged, a fact worth knowing for statement-based questions.

What is the difference between k-fold and leave-one-out cross-validation?

In k-fold cross-validation you split the data into kk folds, train kk models each on k−1k - 1 folds, and average the validation scores. Leave-one-out is the case k=nk = n: it trains nn models, each on n−1n - 1 points. It is nearly unbiased but expensive and has high variance. Never tune on the test set.

How do you compute the variance explained by a principal component?

Centre the data and find the eigenvalues of its covariance matrix. The variance along component kk is its eigenvalue λk\lambda_k, and the share explained is λk∑jλj\displaystyle \frac{\lambda_k}{\sum_j \lambda_j}. With eigenvalues 6, 3 and 1, the first component explains 60 per cent and the first two explain 90 per cent. PCA is sensitive to feature scaling.

Sources

Dates, fees and the syllabus are set by the GATE 2027 organising institute and can change. Always confirm at gate2027.iitm.ac.in.

Keep reading

GATE DA 2027 book614 pages · ₹250 ₹300
Buy now — ₹250