GATE GUIDE

GATE DA Formula Sheet 2027: Key Formulas by Subject

By MD ANISH AHAMADUpdated 4 Oct 20268 min read
GATE DA Formula Sheet 2027: Key Formulas by Subject

This GATE DA formula sheet collects the key formulas of every technical section of the 2027 syllabus, with the trap to watch for beside each one. It is a free selection from the last-minute sheet in the GATE DA 2027 book. Use it to test your recall, then practise each formula on previous-year questions.

In this guide
  1. Key takeaways
  2. How to read this sheet
  3. Probability and Statistics
  4. Linear Algebra
  5. Calculus and Optimisation
  6. Programming, Data Structures and Algorithms
  7. Database Management and Warehousing
  8. Machine Learning
  9. Search, Logic and Reasoning Under Uncertainty
  10. General Aptitude essentials
  11. How to use this sheet in the last weeks
  12. Subject-wise formula sheets
  13. Quick revision

Key takeaways

How to read this sheet

A formula here is a result you apply directly. Its condition is what must be true for it to hold. Write both from memory, or you have learned only half the entry.

Notation follows the book. ln⁡\ln is the natural logarithm; log⁡2\log_2 is written out wherever bits are meant. P(A∣B)P(A \mid B) is the probability of AA given BB. Vectors are columns and x⊤x^\top is the transpose. nn is the number of observations or elements unless stated.

Probability and Statistics

Formula Watch out for
P(A∣B)=P(A∩B)P(B)\displaystyle P(A \mid B) = \frac{P(A \cap B)}{P(B)}, P(B)>0P(B) > 0 P(A∣B)P(A \mid B) is not P(B∣A)P(B \mid A)
P(Ai∣B)=P(B∣Ai) P(Ai)∑jP(B∣Aj) P(Aj)\displaystyle P(A_i \mid B) = \frac{P(B \mid A_i)\,P(A_i)}{\sum_j P(B \mid A_j)\,P(A_j)} The AjA_j must form a partition
Independent: P(A∩B)=P(A) P(B)P(A \cap B) = P(A)\,P(B); mutually exclusive: P(A∩B)=0P(A \cap B) = 0 Events with non-zero probability cannot be both
Var⁡(X)=E[X2]−(E[X])2\operatorname{Var}(X) = E[X^2] - (E[X])^2; Var⁡(aX+b)=a2Var⁡(X)\operatorname{Var}(aX + b) = a^2 \operatorname{Var}(X) a2a^2, not aa; the shift bb drops out
Var⁡(X+Y)=Var⁡(X)+Var⁡(Y)+2Cov⁡(X,Y)\operatorname{Var}(X + Y) = \operatorname{Var}(X) + \operatorname{Var}(Y) + 2\operatorname{Cov}(X, Y) The covariance term vanishes only if uncorrelated
ρ=Cov⁡(X,Y)σXσY\displaystyle \rho = \frac{\operatorname{Cov}(X, Y)}{\sigma_X \sigma_Y}, −1≤ρ≤1-1 \le \rho \le 1 Uncorrelated does not imply independent
s2=1n−1∑(xi−xˉ)2\displaystyle s^2 = \frac{1}{n - 1}\sum (x_i - \bar{x})^2 Divide by nn only for the population form
E[X]=E[E[X∣Y]]E[X] = E\big[E[X \mid Y]\big]; Var⁡(X)=E[Var⁡(X∣Y)]+Var⁡(E[X∣Y])\operatorname{Var}(X) = E[\operatorname{Var}(X \mid Y)] + \operatorname{Var}(E[X \mid Y]) Two terms in the total variance, not one
Distribution Mean Variance
Binomial(n,p)\text{Binomial}(n, p) npnp np(1−p)np(1 - p)
Poisson(λ)\text{Poisson}(\lambda) λ\lambda λ\lambda
Geometric, trial of first success 1p\displaystyle \frac{1}{p} 1−pp2\displaystyle \frac{1 - p}{p^2}
Uniform(a,b)\text{Uniform}(a, b) a+b2\displaystyle \frac{a + b}{2} (b−a)212\displaystyle \frac{(b - a)^2}{12}
Exponential(λ)\text{Exponential}(\lambda) 1λ\displaystyle \frac{1}{\lambda} 1λ2\displaystyle \frac{1}{\lambda^2}
χ2\chi^2 with kk degrees of freedom kk 2k2k

For inference, the central limit theorem gives Xˉ≈N ⁣(μ,σ2n)\displaystyle \bar{X} \approx N\!\left(\mu, \frac{\sigma^2}{n}\right) for independent draws with finite variance. A 95 per cent interval with known σ\sigma is xˉ±1.96 σn\displaystyle \bar{x} \pm 1.96\,\frac{\sigma}{\sqrt{n}}. With σ\sigma unknown, use tt with n−1n - 1 degrees of freedom. The chi-squared statistic is ∑(O−E)2E\displaystyle \sum \frac{(O - E)^2}{E}, and a test of independence has (r−1)(c−1)(r - 1)(c - 1) degrees of freedom.

A one-line check of your own: if Var⁡(X)=4\operatorname{Var}(X) = 4, then Var⁡(3X+2)=9×4=36\operatorname{Var}(3X + 2) = 9 \times 4 = 36.

Trap: The exponential variance is 1/λ21/\lambda^2, not 1/λ1/\lambda. Only the exponential (continuous) and the geometric (discrete) are memoryless.

Linear Algebra

Formula Watch out for
rank⁡A+nullity⁡A=number of columns\operatorname{rank} A + \operatorname{nullity} A = \text{number of columns} Columns, not rows
Ax=bAx = b has a unique solution iff rank⁡A=rank⁡[A ∣ b]=n\operatorname{rank} A = \operatorname{rank}[A \,\vert\, b] = n Equal ranks below nn give infinitely many
∑λi=tr⁡A\sum \lambda_i = \operatorname{tr} A, ∏λi=det⁡A\prod \lambda_i = \det A; 2×22 \times 2: λ2−(tr⁡A)λ+det⁡A=0\lambda^2 - (\operatorname{tr} A)\lambda + \det A = 0 Count repeated eigenvalues
AkA^k has λk\lambda^k; A−1A^{-1} has 1λ\displaystyle \frac{1}{\lambda}; A+cIA + cI has λ+c\lambda + c Eigenvectors stay the same
det⁡(kA)=kndet⁡A\det(kA) = k^n \det A for n×nn \times n Not kdet⁡Ak \det A
P=A(A⊤A)−1A⊤P = A(A^\top A)^{-1}A^\top; P⊤=PP^\top = P, P2=PP^2 = P Needs independent columns
Idempotent: eigenvalues 00 or 11, rank⁡=tr⁡\operatorname{rank} = \operatorname{tr} I−PI - P is idempotent too
Positive definite iff all λ>0\lambda > 0 iff all leading principal minors >0> 0 Leading minors ≥0\ge 0 do not prove semidefinite
A=UΣV⊤A = U\Sigma V^\top; σi2\sigma_i^2 are eigenvalues of A⊤AA^\top A; ∥A∥2=σ1\lVert A \rVert_2 = \sigma_1 Rank is the number of non-zero σi\sigma_i

Calculus and Optimisation

Programming, Data Structures and Algorithms

Formula Watch out for
Binary search worst case: ⌊log⁡2n⌋+1\lfloor \log_2 n \rfloor + 1 probes The array must be sorted
Selection sort: always n(n−1)2\displaystyle \frac{n(n - 1)}{2} comparisons, at most n−1n - 1 swaps Same count on sorted input
Insertion sort: n−1n - 1 comparisons sorted, n(n−1)2\displaystyle \frac{n(n - 1)}{2} reverse-sorted Shifts equal the number of inversions
Merging lists of lengths mm and nn: at most m+n−1m + n - 1 comparisons Not m+nm + n
Quicksort worst: T(n)=T(n−1)+Θ(n)=Θ(n2)T(n) = T(n - 1) + \Theta(n) = \Theta(n^2) Sorted input with an end pivot
Chaining: expected search 1+α1 + \alpha, α=nm\displaystyle \alpha = \frac{n}{m}; open addressing: at most 11−α\displaystyle \frac{1}{1 - \alpha} probes unsuccessful α\alpha is keys over slots
Height hh in edges: h+1≤nodes≤2h+1−1h + 1 \le \text{nodes} \le 2^{h+1} - 1 Check whether height counts edges or levels
Binary tree shapes on nn nodes: 1n+1(2nn)\displaystyle \frac{1}{n + 1}\binom{2n}{n} 1, 2, 5, 14, 42 for n=1n = 1 to 5
Dijkstra O((V+E)log⁡V)O((V + E)\log V); Bellman–Ford O(VE)O(VE); Floyd–Warshall O(V3)O(V^3) Dijkstra needs non-negative weights

Python 3 rules decide answers too: -7 // 2 is -4, -7 % 2 is 1, round(2.5) is 2, and range(a, b) stops before b.

Database Management and Warehousing

This page is a selection. The book's last-minute sheet covers every topic in all seven technical sections, with each result's condition printed beside it. It is part of the GATE DA 2027 book, which also has 907 questions with worked solutions and 10 full mock tests.

Machine Learning

Formula Watch out for
β1=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2\displaystyle \beta_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2}, β0=yˉ−β1xˉ\beta_0 = \bar{y} - \beta_1 \bar{x} The line passes through (xˉ,yˉ)(\bar{x}, \bar{y})
β^=(X⊤X)−1X⊤y\hat{\beta} = (X^\top X)^{-1}X^\top y; ridge: (X⊤X+λI)−1X⊤y(X^\top X + \lambda I)^{-1}X^\top y Ridge shrinks, never to exactly zero
P(y=1∣x)=σ(w⊤x+b)P(y = 1 \mid x) = \sigma(w^\top x + b) The boundary w⊤x+b=0w^\top x + b = 0 is linear
P(C∣x)∝P(C)∏jP(xj∣C)P(C \mid x) \propto P(C)\prod_j P(x_j \mid C) Features independent given the class
SVM margin =2∥w∥\displaystyle = \frac{2}{\lVert w \rVert} Only support vectors fix ww
H=−∑kpklog⁡2pkH = -\sum_k p_k \log_2 p_k; Gini=1−∑kpk2\text{Gini} = 1 - \sum_k p_k^2 Entropy in bits uses log⁡2\log_2
Precision TPTP+FP\displaystyle \frac{TP}{TP + FP}; recall TPTP+FN\displaystyle \frac{TP}{TP + FN}; F1=2PRP+R\displaystyle F_1 = \frac{2PR}{P + R} Do not swap FP and FN
Expected error =bias2+variance+noise= \text{bias}^2 + \text{variance} + \text{noise} Larger kk in k-NN raises bias
PCA share of component kk: λk∑jλj\displaystyle \frac{\lambda_k}{\sum_j \lambda_j} Centre the data first

A worked line of my own: a node with 3 positive and 1 negative example has H=−(0.75log⁡20.75+0.25log⁡20.25)≈0.811H = -(0.75\log_2 0.75 + 0.25\log_2 0.25) \approx 0.811 bits and Gini =1−(0.5625+0.0625)=0.375= 1 - (0.5625 + 0.0625) = 0.375.

Remember: Layers of sizes n0,n1,…,nLn_0, n_1, \ldots, n_L have ∑lnl−1nl\sum_l n_{l-1} n_l weights plus ∑l≥1nl\sum_{l \ge 1} n_l biases. A 10–5–2 network has 60 weights and 67 parameters.

Search, Logic and Reasoning Under Uncertainty

Formula Watch out for
BFS time and space O(bd)O(b^d); DFS time O(bm)O(b^m), space O(bm)O(bm); iterative deepening time O(bd)O(b^d), space O(bd)O(bd) dd is goal depth, mm maximum depth
A*: f(n)=g(n)+h(n)f(n) = g(n) + h(n); admissible h(n)≤h∗(n)h(n) \le h^*(n); consistent h(n)≤c(n,n′)+h(n′)h(n) \le c(n, n') + h(n') Tree search needs admissible, graph search consistent
Alpha-beta: prune when α≥β\alpha \ge \beta; best ordering O(bm/2)O(b^{m/2}) The root value equals minimax
p→q≡¬p∨q≡¬q→¬pp \to q \equiv \neg p \lor q \equiv \neg q \to \neg p The converse is not equivalent
KB⊨αKB \models \alpha iff KB∧¬αKB \land \neg\alpha is unsatisfiable An unsatisfiable KB entails everything
¬∀x P(x)≡∃x ¬P(x)\neg\forall x\, P(x) \equiv \exists x\, \neg P(x) "Some A is B" is ∃x (A(x)∧B(x))\exists x\,(A(x) \land B(x))
P(x1,…,xn)=∏iP(xi∣parents(Xi))P(x_1, \ldots, x_n) = \prod_i P(x_i \mid \text{parents}(X_i)) A Boolean node with kk Boolean parents needs 2k2^k numbers
Likelihood weighting: weight =∏iP(ei∣parents(Ei))= \prod_i P(e_i \mid \text{parents}(E_i)) No finite sample is exact

Trap: In a collider A→B←CA \to B \leftarrow C, AA and CC are independent until BB, or any descendant of BB, is observed. Observing it makes them dependent.

General Aptitude essentials

General Aptitude is worth 15 marks in GATE DA. The DA papers have leaned on number sense rather than commercial arithmetic. The General Aptitude preparation guide covers the whole section.

How to use this sheet in the last weeks

Cover the right-hand column and recall each condition before you check. Mark every entry you miss and work one small example of your own for it. If you also know the CS syllabus, GATE CS vs GATE DA shows where the two papers overlap, so you can see which formulas you already own.

Practise your arithmetic on the GATE virtual calculator, and carry full precision until the final answer. The MCQ, MSQ and NAT marking scheme, common to every GATE paper, has negative marks only for MCQs. Check the exam-day rules before you travel.

In one line: Write the condition next to every formula, and the four classic slips have nowhere to hide.

Subject-wise formula sheets

Each of these goes deeper on one part of the GATE DA paper, with a worked example and the trap for every formula:

Quick revision

  1. Sample variance divides by n−1n - 1, and Var⁡(aX+b)=a2Var⁡(X)\operatorname{Var}(aX + b) = a^2 \operatorname{Var}(X).
  2. Write the conditional the question asks for, P(A∣B)P(A \mid B) or P(B∣A)P(B \mid A), before substituting.
  3. Entropy uses log⁡2\log_2; the logistic loss uses ln⁡\ln.
  4. rank⁡A+nullity⁡A\operatorname{rank} A + \operatorname{nullity} A equals the number of columns.
  5. Gradient descent on ax2ax^2 converges only for 0<η<1a\displaystyle 0 < \eta < \frac{1}{a}.
  6. Dijkstra needs non-negative weights; Bellman–Ford handles negative ones.
  7. A* tree search needs an admissible heuristic; graph search needs a consistent one.
  8. range(a, b) and s[i:j] stop one short of the end.

Frequently asked questions

Can I take a formula sheet into the GATE DA exam?

Plan as if you cannot. Every formula you need should be in memory before the exam, together with the condition under which it holds. The exam gives you an on-screen virtual calculator, but a calculator cannot tell you which expression to use. Confirm the current exam-day rules, including what you may carry, at gate2027.iitm.ac.in.

Should I use log base 2 or the natural log in GATE DA?

It depends on the quantity. Entropy and information gain in bits use log⁡2\log_2. Likelihoods and the logistic regression loss use the natural logarithm ln⁡\ln. The two differ by a factor of ln⁡2≈0.693\ln 2 \approx 0.693, so the wrong base gives a wrong numerical answer even when the method is right. Read which unit the question asks for.

Does sample variance divide by n or n minus 1?

The unbiased sample variance divides by n−1n - 1: s2=1n−1∑(xi−xˉ)2\displaystyle s^2 = \frac{1}{n-1}\sum (x_i - \bar{x})^2. Dividing by nn gives the population form. GATE DA questions can ask for either, so check the wording. Also remember that Var⁡(aX+b)=a2Var⁡(X)\operatorname{Var}(aX + b) = a^2 \operatorname{Var}(X), which is the other variance slip that often costs marks.

What is the closed-form solution of ridge regression?

Ridge regression minimises ∥y−Xβ∥2+λ∥β∥2\lVert y - X\beta \rVert^2 + \lambda \lVert \beta \rVert^2, which gives β^=(X⊤X+λI)−1X⊤y\hat{\beta} = (X^\top X + \lambda I)^{-1} X^\top y. The matrix is invertible for every λ>0\lambda > 0. As λ\lambda grows, the coefficients shrink towards zero but not exactly to zero, bias rises and variance falls. Setting λ=0\lambda = 0 gives ordinary least squares.

When is A* search optimal?

A* tree search is optimal when the heuristic is admissible, meaning h(n)h(n) never exceeds the true cost to the goal. A* graph search needs a consistent heuristic, h(n)≤c(n,n′)+h(n′)h(n) \le c(n, n') + h(n') with h(goal)=0h(\text{goal}) = 0. Every consistent heuristic is admissible. With h=0h = 0, A* becomes uniform-cost search.

Sources

Dates, fees and the syllabus are set by the GATE 2027 organising institute and can change. Always confirm at gate2027.iitm.ac.in.

Keep reading

GATE DA 2027 book614 pages · ₹250 ₹300
Buy now — ₹250