GATE GUIDE

GATE DA Probability and Statistics Formula Sheet with Traps

By MD ANISH AHAMADUpdated 4 Oct 20267 min read
GATE DA Probability and Statistics Formula Sheet with Traps

The probability and statistics formulas you need for GATE DA fit on a few tables: counting and probability rules, expectation and variance, a dozen distributions, and the sampling results behind confidence intervals and tests. This sheet gives each formula with its condition, a one-line example of my own and the trap that costs marks. Every formula is checked against the last-minute sheet in the GATE DA 2027 book.

In this guide
  1. Key takeaways
  2. The terms this sheet uses
  3. Counting and probability rules
  4. Expectation, variance and correlation
  5. Distributions: PMF, mean and variance
  6. Sampling and the central limit theorem
  7. Confidence intervals and tests
  8. Using the sheet in the exam
  9. Quick revision

Key takeaways

The terms this sheet uses

A random variable is a number attached to each outcome of an experiment. A discrete one has a probability mass function (PMF), P(X=k)P(X = k). A continuous one has a probability density function (PDF), f(x)f(x), whose area over an interval is a probability.

The cumulative distribution function (CDF) is F(x)=P(X≤x)F(x) = P(X \le x). It never decreases and runs from 0 to 1. A sample is the data you observe; the population is what it was drawn from. xˉ\bar{x} is the sample mean, ss the sample standard deviation, and μ\mu and σ\sigma their population counterparts. ln⁡\ln is the natural logarithm.

Counting and probability rules

Formula Watch out for
nPr=n!(n−r)!\displaystyle {}^nP_r = \frac{n!}{(n - r)!}, (nr)=n!r! (n−r)!\displaystyle \binom{n}{r} = \frac{n!}{r!\,(n - r)!} Order matters only for permutations
Arrangements with repeats: n!n1! n2!⋯nk!\displaystyle \frac{n!}{n_1!\,n_2!\cdots n_k!} Divide once for each repeated group
Non-negative solutions of x1+⋯+xk=nx_1 + \cdots + x_k = n: (n+k−1k−1)\binom{n + k - 1}{k - 1}; positive: (n−1k−1)\binom{n - 1}{k - 1} Read "non-negative" against "positive"
P(A∪B)=P(A)+P(B)−P(A∩B)P(A \cup B) = P(A) + P(B) - P(A \cap B) Subtract the overlap once
P(at least one)=1−P(none)P(\text{at least one}) = 1 - P(\text{none}) The fastest route for "at least"
P(A∣B)=P(A∩B)P(B)\displaystyle P(A \mid B) = \frac{P(A \cap B)}{P(B)}, P(B)>0P(B) > 0 Not symmetric in AA and BB
P(B)=∑iP(B∣Ai) P(Ai)P(B) = \sum_i P(B \mid A_i)\,P(A_i) The AiA_i must partition the space
P(Ai∣B)=P(B∣Ai) P(Ai)∑jP(B∣Aj) P(Aj)\displaystyle P(A_i \mid B) = \frac{P(B \mid A_i)\,P(A_i)}{\sum_j P(B \mid A_j)\,P(A_j)} The denominator is total probability

Example: x1+x2+x3+x4=6x_1 + x_2 + x_3 + x_4 = 6 has (93)=84\binom{9}{3} = 84 non-negative solutions and (53)=10\binom{5}{3} = 10 positive ones.

Bayes in one line: a condition affects 2 per cent of people, a test detects it 90 per cent of the time and gives a false positive 5 per cent of the time. Then P(ill∣+)=0.0180.018+0.049≈0.269\displaystyle P(\text{ill} \mid +) = \frac{0.018}{0.018 + 0.049} \approx 0.269.

Trap: A test that is 90 per cent accurate does not make a positive result 90 per cent reliable. The small prior drags the posterior down, as the example shows.

Expectation, variance and correlation

Formula Watch out for
E[aX+bY+c]=aE[X]+bE[Y]+cE[aX + bY + c] = aE[X] + bE[Y] + c No independence needed
Var⁡(X)=E[X2]−(E[X])2\operatorname{Var}(X) = E[X^2] - (E[X])^2 Square of the mean, not mean of the square
Var⁡(aX+b)=a2Var⁡(X)\operatorname{Var}(aX + b) = a^2 \operatorname{Var}(X) The shift bb disappears
Var⁡(X+Y)=Var⁡(X)+Var⁡(Y)+2Cov⁡(X,Y)\operatorname{Var}(X + Y) = \operatorname{Var}(X) + \operatorname{Var}(Y) + 2\operatorname{Cov}(X, Y) Drop the last term only if uncorrelated
Cov⁡(X,Y)=E[XY]−E[X]E[Y]\operatorname{Cov}(X, Y) = E[XY] - E[X]E[Y] E[XY]=E[X]E[Y]E[XY] = E[X]E[Y] needs independence or zero covariance
ρ=Cov⁡(X,Y)σXσY\displaystyle \rho = \frac{\operatorname{Cov}(X, Y)}{\sigma_X \sigma_Y}, −1≤ρ≤1-1 \le \rho \le 1 Unchanged by positive rescaling
E[X]=E[E[X∣Y]]E[X] = E\big[E[X \mid Y]\big] The law of total expectation
Var⁡(X)=E[Var⁡(X∣Y)]+Var⁡(E[X∣Y])\operatorname{Var}(X) = E[\operatorname{Var}(X \mid Y)] + \operatorname{Var}(E[X \mid Y]) Both terms, every time
s2=1n−1∑(xi−xˉ)2\displaystyle s^2 = \frac{1}{n - 1}\sum (x_i - \bar{x})^2 Population form divides by nn
New mean after one point: nxˉ+xnewn+1\displaystyle \frac{n\bar{x} + x_{\text{new}}}{n + 1} Recompute the total, then divide

Examples: if Var⁡(X)=5\operatorname{Var}(X) = 5, then Var⁡(2X−3)=20\operatorname{Var}(2X - 3) = 20. If Cov⁡(X,Y)=6\operatorname{Cov}(X, Y) = 6, σX=2\sigma_X = 2 and σY=5\sigma_Y = 5, then ρ=0.6\rho = 0.6. Four values with mean 10 plus a new value 20 give a mean of 40+205=12\displaystyle \frac{40 + 20}{5} = 12.

The median of a continuous variable is the mm with F(m)=0.5F(m) = 0.5. For data, it is the middle sorted value, or the mean of the two middle values when nn is even. The mode is the most frequent value, or the maximiser of a density.

Distributions: PMF, mean and variance

Distribution PMF or PDF Mean Variance
Bernoulli(p)\text{Bernoulli}(p) P(1)=pP(1) = p pp p(1−p)p(1 - p)
Binomial(n,p)\text{Binomial}(n, p) (nk)pk(1−p)n−k\binom{n}{k}p^k(1 - p)^{n - k} npnp np(1−p)np(1 - p)
Uniform on {1,…,n}\{1, \ldots, n\} 1n\displaystyle \frac{1}{n} n+12\displaystyle \frac{n + 1}{2} n2−112\displaystyle \frac{n^2 - 1}{12}
Geometric, k≥1k \ge 1 (1−p)k−1p(1 - p)^{k - 1}p 1p\displaystyle \frac{1}{p} 1−pp2\displaystyle \frac{1 - p}{p^2}
Poisson(λ)\text{Poisson}(\lambda) e−λλkk!\displaystyle \frac{e^{-\lambda}\lambda^k}{k!} λ\lambda λ\lambda
Uniform(a,b)\text{Uniform}(a, b) 1b−a\displaystyle \frac{1}{b - a} a+b2\displaystyle \frac{a + b}{2} (b−a)212\displaystyle \frac{(b - a)^2}{12}
Exponential(λ)\text{Exponential}(\lambda) λe−λx\lambda e^{-\lambda x}, x≥0x \ge 0 1λ\displaystyle \frac{1}{\lambda} 1λ2\displaystyle \frac{1}{\lambda^2}
N(μ,σ2)N(\mu, \sigma^2) 1σ2πe−(x−μ)2/(2σ2)\displaystyle \frac{1}{\sigma\sqrt{2\pi}}e^{-(x - \mu)^2/(2\sigma^2)} μ\mu σ2\sigma^2
χk2\chi^2_k sum of kk squared N(0,1)N(0, 1) kk 2k2k
tνt_\nu symmetric, heavier tails 00 (ν>1\nu > 1) νν−2\displaystyle \frac{\nu}{\nu - 2} (ν>2\nu > 2)

Examples: Binomial(10,0.3)\text{Binomial}(10, 0.3) has mean 3 and variance 2.1. For Exponential(0.5)\text{Exponential}(0.5), P(X>4)=e−2≈0.135P(X > 4) = e^{-2} \approx 0.135 and the median is ln⁡20.5≈1.386\displaystyle \frac{\ln 2}{0.5} \approx 1.386.

Fact Watch out for
P(X>s+t∣X>s)=P(X>t)P(X > s + t \mid X > s) = P(X > t) Only exponential and geometric are memoryless
Binomial(n,p)≈Poisson(np)\text{Binomial}(n, p) \approx \text{Poisson}(np) Needs large nn and small pp
aX+b∼N(aμ+b, a2σ2)aX + b \sim N(a\mu + b,\, a^2\sigma^2); Z=X−μσ\displaystyle Z = \frac{X - \mu}{\sigma} Divide by σ\sigma, not σ2\sigma^2
Φ(1.645)=0.95\Phi(1.645) = 0.95, Φ(1.96)=0.975\Phi(1.96) = 0.975, Φ(2.576)=0.995\Phi(2.576) = 0.995; Φ(−z)=1−Φ(z)\Phi(-z) = 1 - \Phi(z) Two-sided tails split α\alpha in half
P(a<X≤b)=F(b)−F(a)P(a < X \le b) = F(b) - F(a) A discrete CDF jumps at each value

Remember: Independent Poissons add with their rates added. Independent normals add with means added and variances added, never standard deviations.

This sheet covers one section. The book's last-minute sheet does the same for all seven technical sections, each result with its condition beside it. It sits in the GATE DA 2027 book alongside 907 questions with worked solutions and 10 full mock tests.

Sampling and the central limit theorem

The central limit theorem says a sum or mean of many independent draws is close to normal, whatever the original shape, provided the variance is finite.

Formula Watch out for
Sn≈N(nμ, nσ2)S_n \approx N(n\mu,\, n\sigma^2); Xˉ≈N ⁣(μ,σ2n)\displaystyle \bar{X} \approx N\!\left(\mu, \frac{\sigma^2}{n}\right) Finite variance and independence required
Standard error of a mean: σn\displaystyle \frac{\sigma}{\sqrt{n}}; of a proportion: p(1−p)n\displaystyle \sqrt{\frac{p(1 - p)}{n}} n\sqrt{n}, not nn
(n−1)s2σ2∼χn−12\displaystyle \frac{(n - 1)s^2}{\sigma^2} \sim \chi^2_{n - 1}; xˉ−μs/n∼tn−1\displaystyle \frac{\bar{x} - \mu}{s/\sqrt{n}} \sim t_{n - 1} Normal samples only

Example: the sum of 100 independent Bernoulli(0.2)\text{Bernoulli}(0.2) variables has mean 20, variance 16 and standard deviation 4.

Confidence intervals and tests

Formula Watch out for
xˉ±zα/2 σn\displaystyle \bar{x} \pm z_{\alpha/2}\,\frac{\sigma}{\sqrt{n}} σ\sigma known
xˉ±tα/2, n−1 sn\displaystyle \bar{x} \pm t_{\alpha/2,\, n - 1}\,\frac{s}{\sqrt{n}} σ\sigma unknown
p^±zα/2p^(1−p^)n\displaystyle \hat{p} \pm z_{\alpha/2}\sqrt{\frac{\hat{p}(1 - \hat{p})}{n}} Uses the sample proportion
Width halves when nn is multiplied by 4 Not by 2
z=xˉ−μ0σ/n\displaystyle z = \frac{\bar{x} - \mu_0}{\sigma/\sqrt{n}}; reject at 5 per cent two-sided if ∣z∣>1.96\lvert z \rvert > 1.96 One-sided uses 1.645
t=xˉ−μ0s/n\displaystyle t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}}, n−1n - 1 degrees of freedom A paired test is a one-sample test on differences
Pooled two-sample: sp2=(n1−1)s12+(n2−1)s22n1+n2−2\displaystyle s_p^2 = \frac{(n_1 - 1)s_1^2 + (n_2 - 1)s_2^2}{n_1 + n_2 - 2} n1+n2−2n_1 + n_2 - 2 degrees of freedom
χ2=∑(O−E)2E\displaystyle \chi^2 = \sum \frac{(O - E)^2}{E} Degrees of freedom k−1−(parameters estimated)k - 1 - (\text{parameters estimated})
Independence: E=row total×column totalgrand total\displaystyle E = \frac{\text{row total} \times \text{column total}}{\text{grand total}} Degrees of freedom (r−1)(c−1)(r - 1)(c - 1)
Variance test: χ2=(n−1)s2σ02\displaystyle \chi^2 = \frac{(n - 1)s^2}{\sigma_0^2} n−1n - 1 degrees of freedom

Examples: with xˉ=80\bar{x} = 80, σ=10\sigma = 10 and n=25n = 25, the 95 per cent interval is 80±3.9280 \pm 3.92. With xˉ=52\bar{x} = 52, μ0=50\mu_0 = 50, s=6s = 6 and n=9n = 9, t=1t = 1 on 8 degrees of freedom. A 2×32 \times 3 table has 2 degrees of freedom.

A Type I error rejects a true null hypothesis, with probability α\alpha. A Type II error keeps a false one, with probability β\beta; the power is 1−β1 - \beta. The p-value is the probability, under the null, of a statistic at least as extreme as the one observed. Reject when p<αp < \alpha.

In one line: Pick zz or tt by whether σ\sigma is known, count the degrees of freedom, then compare with the critical value.

Using the sheet in the exam

These questions often end in a decimal answer. Practise the arithmetic on the GATE virtual calculator, and round only the final value. Before you guess on a probability MCQ, read the MCQ, MSQ and NAT marking scheme, common to every GATE paper, because only MCQs carry negative marks.

If you know the CS syllabus, GATE CS vs GATE DA shows that statistical inference belongs to DA alone. The General Aptitude guide covers the simpler probability questions in that section.

More formula sheets: all of GATE DA · DBMS and Algorithms · Linear Algebra and Calculus · Machine Learning

Quick revision

  1. Bayes: posterior equals likelihood times prior, over total probability.
  2. Var⁡(aX+b)=a2Var⁡(X)\operatorname{Var}(aX + b) = a^2 \operatorname{Var}(X), and s2s^2 divides by n−1n - 1.
  3. Independence implies zero correlation; zero correlation does not imply independence.
  4. Exponential: mean 1λ\displaystyle \frac{1}{\lambda}, variance 1λ2\displaystyle \frac{1}{\lambda^2}, median ln⁡2λ\displaystyle \frac{\ln 2}{\lambda}.
  5. Poisson mean and variance are both λ\lambda; χk2\chi^2_k has mean kk and variance 2k2k.
  6. Xˉ≈N ⁣(μ,σ2n)\displaystyle \bar{X} \approx N\!\left(\mu, \frac{\sigma^2}{n}\right); quadruple nn to halve the interval width.
  7. 1.96 for 95 per cent two-sided, 1.645 one-sided, 2.576 for 99 per cent.
  8. Chi-squared independence has (r−1)(c−1)(r - 1)(c - 1) degrees of freedom.

Frequently asked questions

What is the variance of an exponential distribution?

For Exponential(λ)\text{Exponential}(\lambda) the mean is 1λ\displaystyle \frac{1}{\lambda} and the variance is 1λ2\displaystyle \frac{1}{\lambda^2}. The tail is P(X>x)=e−λxP(X > x) = e^{-\lambda x} and the median is ln⁡2λ\displaystyle \frac{\ln 2}{\lambda}. A common slip is to give 1λ\displaystyle \frac{1}{\lambda} as the variance. The exponential is also memoryless, a property it shares only with the geometric distribution among the standard distributions.

What is the difference between independent and mutually exclusive events?

Independent events satisfy P(A∩B)=P(A) P(B)P(A \cap B) = P(A)\,P(B), so knowing one tells you nothing about the other. Mutually exclusive events satisfy P(A∩B)=0P(A \cap B) = 0, so one happening rules out the other. Two events that both have non-zero probability can never be both independent and mutually exclusive, which is a frequent multiple-select trap.

Which value of z is used for a 95 per cent confidence interval?

A two-sided 95 per cent interval uses z=1.96z = 1.96, because Φ(1.96)=0.975\Phi(1.96) = 0.975. A one-sided 5 per cent test uses 1.645 and a 99 per cent two-sided interval uses 2.576. When the population standard deviation is unknown, replace zz with the tt value for n−1n - 1 degrees of freedom and use the sample standard deviation ss.

How many degrees of freedom does a chi-squared test have?

A goodness-of-fit test with kk categories has k−1k - 1 degrees of freedom, minus one more for every parameter you estimate from the data. A test of independence on an r×cr \times c table has (r−1)(c−1)(r - 1)(c - 1). A test for one variance uses n−1n - 1. The expected count in a cell is the row total times the column total, divided by the grand total.

Does zero correlation mean two variables are independent?

No. Independence implies zero correlation, but zero correlation does not imply independence. Correlation measures only linear association. If XX is uniform on (−1,1)(-1, 1) and Y=X2Y = X^2, then YY is completely determined by XX, yet their correlation is zero. So a zero correlation only rules out a linear relationship, never every kind of dependence.

Sources

Dates, fees and the syllabus are set by the GATE 2027 organising institute and can change. Always confirm at gate2027.iitm.ac.in.

Keep reading

GATE DA 2027 book614 pages · ₹250 ₹300
Buy now — ₹250