By the end of this chapter you'll be able to…

  • 1Distinguish primary from secondary data, and quantitative from qualitative data, using the defining feature of each rather than surface appearance
  • 2Distinguish discrete from continuous data, and correctly choose between a bar diagram and a histogram based on that distinction
  • 3Name and apply the five components of a well-constructed statistical table, including why the footnote is exam-critical
  • 4Distinguish a frequency polygon from an ogive, and state which of the two is used to read off the median graphically
  • 5Compute a sector's central angle or actual value from a pie chart, and correctly compare percentage shares across two pie charts with different totals
  • 6Compute totals, averages, percentage change, and ratios from a small data table, and correctly identify percentage growth as distinct from absolute growth
💡
Why this chapter matters in UGC NET / JRF
Data Interpretation is the one Paper 1 unit where the bottleneck is speed and reading discipline rather than knowledge — the underlying arithmetic (percentage change, averages, ratios, shares of a total) is school-level, but NTA scores it under time pressure and plants traps specifically in the base of a percentage, a missing units footnote, or a pie chart that silently changed its total between two years. Because Paper 1 carries a flat +2 per correct answer with zero negative marking, a candidate who reads a table correctly but computes 20 seconds slower loses nothing on marking, but loses real minutes that Reading Comprehension or Logical Reasoning questions elsewhere on the same paper needed. This chapter is deliberately built around worked numeric drills, not just definitions, because the only way to get faster at DI is repeated, timed practice on small data sets until the base-of-a-percentage habit becomes automatic.

Data Interpretation — UGC NET Paper 1

Every other unit on this paper rewards knowing a fact. This one rewards reading carefully. A Data Interpretation question rarely hides a hard concept — it hides a small trap in the base of a percentage, the units line of a table, or a pie chart that quietly changed its total between two years. Candidates who "know" percentages and ratios cold from school still lose these marks, purely to hurry.


1. What UGC NET actually asks

Data Interpretation carries weightPct 9 of Paper 1's 50 questions — roughly 4 to 5 questions, each worth a flat 2 marks with no negative marking, so every question here is worth attempting once you've narrowed it to two options, since a blank answer and a wrong guess score identically.

The unit shows up in two distinct flavours, and NTA mixes both across a sitting:

  1. Conceptual questions with no numbers at all — "which chart is appropriate for this kind of data," "what distinguishes a histogram from a bar diagram," "which of these is an example of secondary data." These are pure recall and reward knowing the vocabulary precisely.
  2. Computational sets — a small table or a described bar/pie/line chart, followed by one or more questions asking for a total, an average, a percentage change, a ratio, or a share of the whole. These reward arithmetic fluency and reading discipline far more than any advanced technique.

Across both flavours, the syllabus itself groups into four connected ideas: sources and classification of data, tabulation, graphical representation, and interpretation exercises built on a given table or graph. The rest of this chapter works through each in turn, then drills the computation with two full worked data sets.


2. Sources and classification of data

Before any chart is drawn, data has to be collected and sorted — and NTA tests the vocabulary of that first step directly.

Primary data is collected firsthand, by the researcher, for the specific purpose at hand — through surveys, questionnaires, interviews, direct observation, or experiments. It is original and precisely tailored to the question being asked, but it costs more time, effort, and money to gather.

Secondary data is data that already exists, collected earlier by someone else for a different original purpose, and reused — government census reports, published journal articles, institutional records, or data available from bodies such as the Ministry of Statistics and Programme Implementation. It is cheaper and faster to access, but it demands a check on the source's reliability, the period it covers, and whether it actually fits the current question, since it wasn't designed for it.

A second, independent classification cuts across this one:

  • Quantitative data is inherently numerical and measurable — marks scored, income earned, population counted, temperature recorded.
  • Qualitative data is categorical or descriptive rather than naturally numeric — a respondent's opinion, a quality rating, a yes/no answer, a colour or category label. It can be coded into numbers for analysis, but the underlying property being recorded is not itself a quantity.

A further, commonly tested split applies specifically within quantitative data:

  • Discrete data takes only countable, whole-number values with no meaningful values in between — the number of students in a class, the number of research papers published, the number of books on a shelf.
  • Continuous data can take any value within a range, including fractional values — height, weight, time taken, temperature.

This last distinction matters practically: it decides whether a bar diagram or a histogram is the correct chart for a given data set, which is exactly the kind of question NTA likes to ask directly.


3. Tabulation of data

Tabulation is the systematic arrangement of raw, unsorted data into rows and columns so it can be read and analysed at a glance, and it is treated as a distinct step, before any chart is drawn. A well-constructed statistical table conventionally carries five parts:

ComponentRole
Table numberIdentifies the table for reference in the surrounding text
TitleStates what the table shows, as precisely as possible
CaptionThe column headings, describing what each column records
StubThe row headings, describing what each row records
BodyThe actual numerical data, at the intersection of rows and columns
Footnote / source noteClarifies units ("figures in lakhs"), exceptions, or the data's origin

The footnote is the single most exam-relevant part of this list: a large share of DI errors trace back to missing a units line ("production in thousand tonnes") and computing an answer off by a factor of a thousand.


4. Graphical representation of data

Once tabulated, data is commonly displayed graphically for faster visual comprehension. NTA's syllabus names several standard forms, and knowing which form suits which kind of data is tested directly.

  • Bar diagram — rectangular bars whose length is proportional to the value represented, with a gap between adjacent bars since the underlying categories are discrete and unordered. Variants include the simple bar diagram (one variable per category), the multiple/grouped bar diagram (two or more variables compared side by side per category — e.g., a college's male and female enrolment shown as paired bars for each year), and the sub-divided (component) bar diagram, where each bar is stacked into segments representing parts of a whole, useful for showing both a total and its composition in a single bar.
  • Histogram — used specifically for continuous grouped data (a frequency distribution over class intervals), drawn with no gaps between adjacent bars, since the class intervals themselves are continuous. With equal class widths, the height of each bar is proportional to its frequency; with unequal widths, it is the area of the bar, not the height, that correctly represents frequency. This is the most commonly tested bar-vs-histogram distinction: a bar diagram compares discrete, unrelated categories with gaps between bars; a histogram displays a continuous distribution with no gaps.
  • Frequency polygon — formed by joining the midpoints of the tops of histogram bars with straight lines, and closed to the horizontal axis at both ends. It is especially useful for comparing two or more frequency distributions on the same axes, which is harder to do cleanly with overlapping histograms.
  • Ogive (cumulative frequency curve) — plots cumulative frequency against the upper (or lower) class boundary, and is the standard graphical route to reading off the median and other percentile values directly from a chart, without needing to compute them from raw data. NTA sometimes tests the ogive purely by name, so don't confuse it with the frequency polygon: a frequency polygon plots frequency itself; an ogive plots cumulative frequency.
  • Pie chart (pie diagram) — a circle divided into sectors, where each sector's central angle is proportional to the share of the total it represents. The controlling formula is:

  • Line graph — plots values against a continuous variable (almost always time), joined by line segments, and is the standard choice for showing a trend — growth, decline, or fluctuation — over successive periods, rather than a one-time comparison across categories.

5. Worked table — reading, totals, averages, and percentage change

Here is a small table of the kind NTA attaches two to four questions to. It shows the number of research papers published by four departments of a university over five years:

Department20202021202220232024
Physics2228253035
Chemistry1820242628
Botany1514182019
Zoology1012111518

Total papers published across all four departments in 2023: read the 2023 column and add — .

Average annual output of the Chemistry department across the five years: sum the row and divide by the number of years — papers per year.

Percentage increase in Physics's output from 2020 to 2024: .

Which department showed the greatest percentage growth from 2020 to 2024 — not the greatest absolute growth? Compute each department's percentage change on its own 2020 base:

  • Physics:
  • Chemistry:
  • Botany:
  • Zoology:

Zoology wins, at 80% — despite having the smallest raw numbers in the table throughout. This is the single most exam-relevant lesson a DI table can teach: the department with the biggest bars is not automatically the department with the biggest percentage growth, because percentage change divides by each series' own starting value, and a smaller base inflates the percentage even for a modest absolute gain.


6. Worked pie chart — degrees, shares, and amounts

A university's annual budget of ₹100 crore is represented as a pie chart with the following sector shares: Salaries 45%, Infrastructure 20%, Research grants 15%, Scholarships 10%, Administration 10%.

Central angle of the Research grants sector: .

Actual amount allocated to Infrastructure: .

Combined share of Scholarships and Administration, as a central angle: .

The two-pie trap, worth flagging even without a second chart here: if a question ever shows the same university's budget as two separate pie charts for two different years, a sector's percentage share can fall while its actual rupee amount rises, if the total budget itself grew between the two years. Percentage shares only compare cleanly within one pie; comparing actual values across two pies requires converting each share back to an amount using that year's own total first.


7. Solved PYQ-style examples

Q1. A researcher directly conducts a structured interview with 200 government school teachers to study their attitude toward continuous assessment. This data is best classified as: Solution. The researcher collected this data firsthand, for this specific study, through direct interviews — the defining feature of primary data. Answer: Primary data.

Q2. A continuous frequency distribution of the marks scored by 60 students, grouped into class intervals of 10, is best represented graphically by a: Solution. Marks grouped into continuous class intervals call for a histogram, drawn with no gaps between adjacent bars, rather than a bar diagram, which is used for discrete, unrelated categories with gaps between bars. Answer: Histogram.

Q3. In a pie chart of a household's monthly expenditure, the "Education" sector occupies a central angle of 72°. If the household's total monthly expenditure is ₹40,000, the amount spent on education is: Solution. Convert the angle to a share of the whole, then apply it to the total: , and . Answer: ₹8,000.

Q4. Using the department table in Section 5, Botany's output fell between which two consecutive years before recovering? Solution. Reading the Botany row (15, 14, 18, 20, 19) year by year, output dropped from 2020 (15) to 2021 (14), then rose continuously through 2023 (20), before dipping again in 2024 (19). The first fall is between 2020 and 2021. Answer: 2020 to 2021.

Q5. A graph is required to compare two frequency distributions of examination marks on the same set of axes, so their shapes can be visually compared without overlapping bars. The most suitable choice is a: Solution. A frequency polygon, formed by joining the midpoints of histogram bar-tops with straight lines, is specifically suited to overlaying and comparing two or more distributions on shared axes, which overlapping histogram bars would make difficult to read. Answer: Frequency polygon.

Q6. To read off the median value of a large, grouped data set directly from a graph without recomputing it from the raw frequencies, the correct graphical tool is a(n): Solution. The ogive plots cumulative frequency against class boundary and is the standard graphical route to locating the median (and other percentiles) directly, distinguishing it from a frequency polygon, which plots frequency itself rather than cumulative frequency. Answer: Ogive (cumulative frequency curve).


8. Common traps

  • Computing percentage change on the wrong base — "increase from A to B" always divides by A (the original/earlier value), never by B; reversing this silently produces a plausible-looking but wrong figure.
  • Assuming the biggest bar means the biggest percentage growth — as Section 5 showed, a series with small absolute values can post the highest percentage growth precisely because its starting base is small; always compute the ratio, don't eyeball the bars.
  • Missing the units footnote — a table or chart marked "in thousands" or "in lakh" changes every computed answer by that factor; scan the footnote before calculating anything.
  • Comparing sector percentages across two different pie charts as if they were on the same total — a falling percentage share can still mean a rising actual amount if the underlying total grew; convert both pies' shares back to actual values before comparing them.
  • Confusing a bar diagram with a histogram — a bar diagram has gaps and compares discrete, unrelated categories; a histogram has no gaps and displays a single continuous frequency distribution across class intervals.
  • Confusing a frequency polygon with an ogive — a frequency polygon plots frequency against class midpoint; an ogive plots cumulative frequency against class boundary, and only the ogive is used to read off the median graphically.
  • Treating secondary data as automatically reliable — because it wasn't collected for the current purpose, its period, method, and source must still be checked before it's used to answer a new question.
  • Rushing the arithmetic instead of budgeting precision to the option spread — if the four answer options are far apart, a fast approximation is safe; if they sit close together, the full exact calculation is unavoidable, so glance at the options before deciding how carefully to compute.

9. Training protocol

Treat this unit as an arithmetic-speed drill, not a concepts unit — the vocabulary (primary/secondary, discrete/continuous, bar/histogram/polygon/ogive/pie) can be fixed solidly in a single sitting, and after that, the only way to actually get faster is repeated timed practice on small tables and pie charts until percentage change, average, ratio, and share-of-total become near-automatic. Always read a data set's title, column headings, and footnote before touching a single number — most of the traps in Section 8 are caught at this reading stage, not the calculating stage. When a question asks for "percentage growth" or "percentage change," write down which value is the base before dividing, every single time, until it becomes a reflex; this one habit alone eliminates the most common wrong answer NTA plants in this unit. And because Paper 1 carries zero negative marking, never leave a Data Interpretation question blank once you've read the chart correctly and eliminated even one option — the arithmetic risk of a small error is still better than the guaranteed zero of skipping it.

Key formulas & results

Everything to memorise for the exam hall, in one card. Screenshot this for revision.

Percentage change
% change = [(New value − Original value) / Original value] × 100
The denominator is always the earlier/original value, never the later one — reversing this is the single most common DI error.
Average (mean) of a data series
Average = (Sum of all values) / (Number of values)
Used for a table's row or column total spread evenly across its count.
Pie chart sector angle and value
Sector angle = (Category value / Total value) × 360°; Category value = (Sector angle / 360°) × Total value
Percentage shares only compare directly within one pie chart; comparing actual values across two different pies requires converting each share back to a value using that pie's own total.
Primary vs secondary data
Primary data = collected firsthand by the researcher for the specific study at hand; Secondary data = pre-existing data collected earlier, for a different purpose, and reused
Primary data is original and tailored but costlier; secondary data is faster to access but must be checked for reliability, period, and fit to the current question.
Discrete vs continuous data
Discrete data = countable, whole-number values only (e.g., number of students); Continuous data = any value within a range, including fractions (e.g., time, height, weight)
This distinction directly decides bar diagram (discrete, gaps between bars) versus histogram (continuous, no gaps).
Frequency polygon vs ogive
Frequency polygon = line joining midpoints of histogram bar-tops, plotting frequency; Ogive = cumulative frequency plotted against class boundary
Only the ogive is used to read off the median and other percentiles directly from a graph.
⚠️

Traps UGC NET / JRF sets — and how to dodge them

These are the exact option-traps and misreads that cost marks under negative marking.

WATCH OUT
Computing percentage change using the later value as the denominator
Always divide by the original/earlier value — 'increase from A to B' means (B−A)/A × 100, with A as the base, every time.
WATCH OUT
Assuming the series with the largest bars also has the largest percentage growth
Compute each series' own percentage change on its own starting value — a small-base series can post the highest percentage growth despite the smallest absolute numbers.
WATCH OUT
Skipping a table or chart's units footnote
Scan the title, headings, and footnote before calculating anything — 'in thousands' or 'in lakh' changes every subsequent answer by that factor.
WATCH OUT
Comparing percentage shares across two different pie charts as if the totals were the same
Convert each pie's percentage share back into an actual value using that specific pie's own total before comparing amounts across two charts.
WATCH OUT
Using a bar diagram's logic for continuous grouped data, or a histogram's logic for discrete categories
Discrete, unrelated categories take a bar diagram with gaps; a single continuous frequency distribution over class intervals takes a histogram with no gaps.
WATCH OUT
Confusing a frequency polygon with an ogive
A frequency polygon plots frequency itself and is best for comparing distribution shapes; an ogive plots cumulative frequency and is the tool for reading off the median.
WATCH OUT
Treating secondary data as automatically trustworthy because it's already published
Check the source, collection period, and whether the data was ever designed to answer the current question before relying on it.

Exam-pattern practice

PYQ-style questions with full solutions. Work through them as a readiness check — mark yourself honestly and get your gap report at the end.

Readiness check

Are you exam-ready for "Data Interpretation"?

12 problems from this chapter. Try each one, reveal the worked solution, mark yourself honestly — get your gap report at the end.

12 questions~8 min
Header Logo