18.3 Rubrics, Portfolios, Lab Journals, and Descriptive Feedback
Key Takeaways
- Product-based assessment is what makes inquiry assessable: a lab journal shows the reasoning process that a multiple-choice item cannot reach.
- A rubric must be shared or co-constructed with students before they begin the work, because criteria withheld until grading measure guessing rather than learning.
- An analytic rubric scores separate dimensions and supports targeted feedback, while a holistic rubric assigns one overall score and is faster but less instructive.
- Descriptive feedback that names the gap and the next step changes performance; a score or a grade alone provides no information the student can act on.
- Portfolios and student profiles document growth over time and support student self-assessment, which is the practice most closely tied to improved achievement in the assessment literature.
Assessing the Process, Not Just the Answer
Competency 023 lists specific instruments — "projects, lab journals, rubrics, portfolios, student profiles, checklists" — and specifies their purpose: evaluating student participation in and understanding of the inquiry process. That purpose is why a science teacher cannot assess inquiry with a multiple-choice test alone. A selected-response item can confirm that a student knows what a controlled variable is; only a product can show whether the student controlled one.
The Instrument Toolkit
| Instrument | What it reveals | Best used for | Limitation |
|---|---|---|---|
| Lab journal / science notebook | Day-to-day thinking: questions, predictions, raw data, revisions, dead ends | Formative monitoring of the inquiry process | Time-consuming to read; needs a focused review protocol |
| Project | Whether a student can carry a multi-step inquiry from question to communication | Summative assessment of integrated skill | Group work can mask individual contribution |
| Portfolio | Growth over time; student reflection on their own progress | Longitudinal evidence; conferences | Requires sustained curation |
| Rubric | Explicit criteria applied consistently across students and time | Both formative and summative scoring | Poorly written descriptors reintroduce subjectivity |
| Checklist | Whether required components are present | Quick completeness checks (data table present? axes labeled?) | Records presence, not quality |
| Student profile | An individual's growth against standards | Conferences, intervention decisions, parent communication | Only as good as the underlying evidence |
| Performance task | Skill demonstrated in real time (using a balance, building a circuit) | Lab-skill assessment | Logistically demanding for a full class |
| Self- and peer assessment | Student ownership of criteria; metacognition | Building independence | Needs modeling and a clear rubric to be reliable |
Making Lab Journals Assessable
A science notebook becomes an assessment instrument only when its structure is defined and taught. A workable grades 4-8 entry structure is: investigative question → prediction with reasoning → procedure → data table → graph → claim-evidence-reasoning conclusion → what I would change. With that structure, a teacher can review the class set for one dimension at a time — this week, whether conclusions cite specific data with units — and turn a two-hour reading task into a fifteen-minute diagnostic sweep.
The "what I would change" line is disproportionately valuable, because it is where students reveal their understanding of experimental error, sample size, and control of variables without being asked a direct question about those terms.
Designing a Rubric
| Feature | Analytic rubric | Holistic rubric |
|---|---|---|
| Structure | Separate criteria, each scored | One overall score |
| Feedback value | High — shows exactly which dimension fell short | Low — one number for the whole product |
| Scoring speed | Slower | Faster |
| Best for | Formative use; complex multi-dimensional products | Quick summative sorting; large volumes |
A usable analytic rubric for an inquiry report might score four dimensions: experimental design (controlled variables, trials), data quality (organization, units, precision), analysis (appropriate graph, accurate reading of the trend), and communication (claim, evidence, reasoning). Each dimension needs descriptive performance levels, not evaluative labels. "Conclusion states a claim, cites at least two specific data values with units, and explains the underlying science" is usable by a student; "excellent conclusion" is not.
Two rubric design errors are common enough to appear as exam distractors: counting instead of judging (awarding points for number of sentences rather than quality of reasoning), and hiding the rubric until grading, which converts an instructional tool into a scoring device.
Sharing Criteria Before the Work
The framework treats sharing evaluation criteria as an expectation, not a courtesy. Two practices satisfy it:
- Distribute or co-construct the rubric before students begin. Students who see the criteria in advance produce measurably stronger work because they can aim at the target instead of inferring it. Co-construction goes further: having a class examine two anonymous sample reports and generate the criteria that distinguish them builds an understanding of quality that no handout transmits.
- Show exemplars. A strong and a weak example, discussed against the rubric, makes abstract descriptors concrete.
Feedback That Changes Performance
A score alone teaches nothing about what to do next. Descriptive feedback identifies the gap between the current work and the criterion and names an actionable next step:
- Not usable: "78% — needs work."
- Usable: "Your claim answers the investigative question clearly. Your evidence does not yet include units or the number of trials — add both to the second paragraph, then explain why three trials makes the pattern more trustworthy."
Effective feedback is task-specific, forward-looking, and timely, and it addresses the work rather than the learner. It also needs an opportunity to be used: feedback delivered on a final graded product that will never be revised changes nothing. Building one revision cycle into a major inquiry report is the structural change that makes feedback function.
A related finding worth knowing: when a score and written comments are returned together, students attend to the score and largely ignore the comments. Returning comments first, with the score after revision, keeps the feedback in play.
Self-Assessment and Student Profiles
Students who assess their own work against a rubric before submitting it internalize the criteria and catch their own omissions. Structured self-assessment — "circle the level you think you reached on each row and highlight the evidence in your report" — takes five minutes and converts the rubric from a teacher tool into a student tool.
A student profile compiles evidence across instruments and time to show growth against specific standards. It is what makes a parent conference concrete ("in September her conclusions restated the hypothesis; by January she was citing data with units and identifying error sources") and what makes an intervention decision defensible.
Alternative Modes Without Lowering the Standard
Emergent bilingual students and students with learning differences can demonstrate science understanding through oral claim-evidence-reasoning responses, labeled diagrams, physical models, recorded explanations, or concept maps. The mode of demonstration changes; the standard does not — the student must still support a claim with evidence and reasoning. Reducing the language load of the directions themselves, through rephrasing, a bilingual glossary, or reading aloud, ensures the assessment measures science understanding rather than reading comprehension.
A teacher grades inquiry reports with a detailed analytic rubric but distributes it only when returning the scored work. What is the primary problem?
Which feedback comment is most likely to improve a student's next inquiry report?
A teacher wants evidence of how students reason during an investigation, including the predictions they revised and the errors they noticed. Which instrument is best suited to that purpose?
You've completed this section
Continue exploring other exams