Founder
September 14, 2026
26 min read
The fact that a grade appears on screen under a teacher’s name does not prove that the teacher actually gave it. If artificial intelligence parsed the response, effectively weighted the criteria, suggested the score, and hundreds of results were approved in a single operation, the real legal question is this: Did the human merely lend their name to the decision, or did they actually form it, explain it, and change it when necessary?
A student’s open-ended examination response is assessed by an artificial intelligence-supported system. The system suggests 72 points. During a busy examination week the teacher reviews results in bulk and approves the suggested scores. The student asks why the method used in their answer was deemed insufficient. The institution has only the final score; the model output at the time of assessment was not retained, the version of the scoring key used could not be identified, and which elements the teacher separately reviewed was not recorded. When the model is run again on the same response it produces 78 points.
In this example the problem is not only “did artificial intelligence make a mistake?” Three questions must be answered first:
To whom will the assessment that produced the grade be legally and institutionally attributed?
What justification will be given to the student showing why their own response led to this outcome?
How will appeal become a genuine re-examination rather than merely reproducing the same system’s result?
The Ministry of National Education Ethical Declaration System for Artificial Intelligence Applications (YAZEK), which opened for use on 2 February 2026, moved the use of artificial intelligence in education beyond a mere technical tool choice into a matter of institutional responsibility, record-keeping, and oversight.[1] The most important message of the system and the accompanying Ethical Guide for Artificial Intelligence Applications in Education on grading is clear: Artificial intelligence may assist; but the ultimate responsibility for outcomes such as passing a class, grades, discipline, and similar results rests with teachers and administrators. Humans must be able to correct the output; the decision process must be recorded and subject to oversight; and a mechanism of appeal must exist for the data subject.[2]
Yet the sentence “ultimate responsibility lies with humans” does not come to life by adding a confirmation button to the screen. That sentence becomes a legal safeguard only when the manner in which assessment authority is exercised, the justification presented to the student, the decision trail, and the route to correction are designed together. Grading compliance with artificial intelligence therefore does not arise from the technology unit testing the model; it arises from measurement and assessment, education law, personal data protection, children’s rights, accessibility, procurement, and internal appeal processes meeting within the same architecture.
This article does not revisit artificial intelligence building student profiles or renegotiating vendor contracts. The focus is narrower, but the consequences more direct: When artificial intelligence contributes to an educational decision, how human oversight will be proven, at which layers justification will be built, and by which procedure appeal will be made effective.
A grade carries at least three distinct qualities at once.
First, a grade is a claim of measurement and assessment. Because it states the extent to which a student has reached a particular learning outcome, it is tested against criteria belonging to educational science such as validity, reliability, objectivity, and usability. MEB regulations on written and practical examinations also require that assessment instruments test the competencies in the curriculum, that results be shown to the student, that feedback be given for incomplete learning, and that an appeal route be operated for certain examinations.[3] Using artificial intelligence does not remove these measurement obligations; on the contrary, it increases the need to make visible by which criterion and evidence the score was formed.
Second, a grade is an institutional exercise of authority. It may affect passing a course, graduation, scholarships, level grouping, programme admission, or disciplinary process. The legal regime of the decision may differ in public school, private school, university, or course context; but the common point is that the outcome is owned by an authorised person or body on behalf of the institution. A service provider’s model or an off-campus platform cannot on its own take over the teacher’s or authorised body’s power to assess.
Third, a grade is the result of processing personal data relating to an identified or identifiable student. Answer text, audio recording, behavioural data, score, feedback, and indicators derived from these may bring the personal data regime into play. For this reason lawfulness, specific and legitimate purpose, data minimisation, accuracy, retention period, security, disclosure, and data-subject rights must be observed not only in model training but also in producing, recording, sharing, and correcting the score.[4]
These three faces do not substitute for one another. A model that appears statistically consistent does not legitimise an unauthorised decision process. An authorised teacher’s signature does not make an invalid assessment instrument scientifically reliable from an educational standpoint. A disclosure text compliant with personal data protection does not provide the student with an academic appeal route by which to correct a wrong grade. Sound compliance design addresses all three together.
Although human responsibility, traceability, and appeal appear as separate expressions in YAZEK documents, for grading they are parts of a single procedural chain:
Human responsibility requires the authorised person to be in a position to understand and change the model output;
Traceability requires that it be possible to show afterwards on which data, criteria, model output, and human intervention the assessment was based;
Appeal requires that a wrong or unfair outcome be effectively re-examined and corrected within the institution.
If one of these elements is missing, the others largely cease to function. Without a record, what the human oversaw cannot be proven. If the student is not given an understandable justification, it becomes unclear what to appeal against. Without authorised and independent re-examination, justification remains merely a one-sided notification. If the human only approves the model result, the presence of a human name in the records does not prove that the decision was materially made by a human.
For this reason the ethical declaration made to YAZEK is not the end of the process but its beginning. The counterpart of the declared principle in daily practice must be visible in job descriptions, screen permissions, scoring workflow, record policy, and appeal file. The real evidence of compliance is not the sentence “human oversight exists” but the decision trail that can show how that oversight was carried out for a particular student.
It is misleading to treat artificial intelligence-supported grading as a single category. The system may only flag spelling errors; group answers before the teacher; suggest draft feedback or scores according to a rubric; or automatically transfer the score to the student information system. The legal value of human intervention is determined not by the tool’s name but by the extent to which it actually directed the outcome.
In an assistive tool, the system edits text, groups similar answers, or marks certain points; the teacher directly assesses the actual response and criterion. In a recommendation system, draft scores or feedback are produced; the teacher must independently evaluate the suggestion and have authority and adequate time to change it. In a de facto determining system, default scores are accepted in bulk or deviation becomes the exception; nominal approval is insufficient—genuine review and intervention trace per student is required, not sampling. In fully automated decision, the outcome is applied directly without meaningful human assessment; the legal basis for use is separately evaluated, and in adverse outcomes data-subject rights including KVKK Art. 11/1-g come into play.
Article 11(1)(g) of Law No. 6698 on the Protection of Personal Data grants the right to object to a result arising to the detriment of the person through analysis of processed data exclusively by automated systems. It cannot be said that this provision automatically applies to every artificial intelligence-supported grade: The conditions of automated analysis, exclusive automatism, and adverse outcome must each be assessed in the concrete case. If there is genuine and effective review by the teacher, the process as a whole may not be regarded as automated. Conversely, it cannot be assumed in advance that a merely formal click will in every case create a “human decision”.[5]
In the Court of Justice of the European Union’s SCHUFA judgment, where a score plays a determining role in a third party’s decision, the production of the score itself was treated as a process that may carry the character of an automated decision.[6] This judgment is neither directly binding under Turkish law nor concerned with educational grading. It nevertheless contains an important comparative warning: Legal characterisation must look not at whose name appears on the final screen of the process but at the extent to which the score actually determined the subsequent decision.
Educational institutions should therefore not rely only on a term of use stating “the final decision rests with the teacher”. The following facts must be examined together:
Does the teacher see the student’s original answer or only the model’s summary?
Are the scoring criteria and the model’s limitations presented in an understandable way?
Does the interface strengthen automation bias by making the suggested score the default?
Does the teacher have enough time to review each student?
Does departing from the system trigger additional approval, explanation, or performance pressure?
When the teacher changes the score, does the change actually reflect in the outcome system?
Can the institution track differences between model suggestion and final grade and human intervention?
The answer to these questions shows whether artificial intelligence is a tool, a recommendation, or the de facto decision-maker.
Human oversight has previously been discussed in educational technology along the axes of “knowledge, authority, and justification”.[7] In grading, the next step needed is to turn this concept into a provable decision standard. For this article that may be called justified human ownership. This expression is not a newly defined right in law or an official YAZEK term; it is a policy criterion that makes existing educational, data protection, and institutional accountability obligations practicable.
For justified human ownership the authorised assessor must:
have access to the student’s actual answer, the scoring criterion in force, and necessary context;
know what function artificial intelligence performed and its main error patterns;
have genuine authority and adequate time to accept, reject, or change the suggestion;
be able to form their own academic assessment independently of the model result;
be able to explain the reason specific to the student for the final outcome;
be able to record appropriately the intervention made or why the suggestion was adopted.
The aim here is not to force the teacher to write a long legal opinion for every score. The aim is to make the answer “the system gave it that way” impossible when the outcome is disputed. The teacher or body must have formed their own professional judgment; and the institution must provide working conditions in which that judgment can actually be exercised.
This criterion also prevents responsibility from being placed on the teacher alone. Inadequate training, excessive workload, an interface that permits only bulk approval, model outputs that are not retained, or version changes not disclosed by the vendor are matters of institutional design and organisation. The principle “ultimate responsibility rests with the teacher” cannot be used as a waiver that eliminates institutional responsibility.
When a student asks “why did I get 72?” three different explanation needs may be conflated.
The first is system explanation regarding how the system works in general: In which courses and tasks is artificial intelligence used, what type of data does it process, does it suggest scores or only produce feedback, what limitations are known, and what is the human’s role?
The second is decision justification relating to a particular student: By which learning outcome and scoring criterion was this answer assessed, which element was found sufficient or insufficient, to what extent did artificial intelligence contribute to the outcome, and what assessment did the authorised human make?
The third is a technical and institutional review file that may need to be opened only to authorised overseers: Details such as model and version used, input-output records, test results, threshold values, access logs, changes, and human intervention are found here.
These layers differ in purpose and level of access.
Advance disclosure is directed at the student and parent; minimum content covers artificial intelligence’s task, data categories, the human’s role, basic limitations, and the application route. Student-specific justification goes to the addressee of the outcome; it should include the rubric applied, relevant answer elements, reasons for the score, artificial intelligence’s contribution, the final decision-maker, and the appeal procedure. Oversight file is for the authorised teacher, appeal authority, data controller, and where necessary overseer or judicial body; it holds the actual work, rubric version, input and output actually used, model/version/time information, human actions, and test and incident records.
Meaningful justification does not mean giving the student the model’s entire mathematical structure or source code. In Dun & Bradstreet Austria the Court of Justice of the European Union emphasised that information on automated decision logic must be suitable for the data subject to understand the decision and object to it—concise and understandable—and that merely transmitting a complex formula may not serve this purpose.[8] The judgment is again not directly binding for Turkish education law. It nevertheless offers a useful distinction for justification design: Technical explainability and the justification of the concrete decision are not the same thing.
Student-specific justification should at minimum answer:
Which question, learning outcome, and rubric version were applied?
Which concrete element in the student’s answer gained or lost points?
What function did artificial intelligence perform: classification, error marking, draft score, similarity detection, or feedback?
Was the artificial intelligence suggestion the same as the final grade; did a human make a correction?
Which authorised person or unit gave the final decision?
Within what period and through which channel can the student request re-examination?
Trade secrets, system security, and third parties’ data must of course be protected. But these interests cannot become the general reason for giving no concrete explanation. Conflicting interests can be balanced by giving the student a functional explanation while keeping more sensitive technical material in an oversight file with restricted access. For effective appeal the student does not need to know the model’s hidden chain of reasoning; they need to know by which criterion their own answer was linked to that outcome.
Generative artificial intelligence systems may give different outputs to the same input at different times. The model or system instruction may change; the vendor may make a silent update; settings such as temperature or external data sources may affect the result. For this reason rerunning the model during appeal does not prove how the first decision was formed.
The first task of appeal should not be to call the model again but to freeze the origin of the decision. A proportionate decision trail should link:
the original work the student submitted and the time of submission;
the question, learning outcome, answer key, or rubric version in force at that moment;
the content actually given to artificial intelligence and the output the system actually produced;
model, provider, version, configuration, and time information to the extent accessible;
which screens the teacher saw, which changes they made, and when they approved the final outcome;
the score, brief justification, and application information communicated to the student;
the decision on appeal and reflection of the correction in all related systems.
Here “retain everything indefinitely” is not the right answer. The scope of recording should be determined by purpose and risk level; student data should not be duplicated unnecessarily; access rights should be limited; retention period should be defined taking into account examination and appeal periods, the institution’s legal obligations, and possible disputes. The integrity, timestamp, access security, and deletion procedure of the decision trail matter as much as its existence.
The European Union Artificial Intelligence Act also, within its scope of application, links human oversight in high-risk systems to the ability to understand the system’s capacity and limitations, recognise automation bias, interpret output, override, reverse, or stop the system. Access to educational institutions, assessment of learning outcomes, and monitoring of prohibited conduct in examinations are counted among high-risk use areas under certain conditions.[9] It cannot be assumed that the Regulation will apply directly to every educational institution and every use in Turkey; national and material scope must be examined separately. It nevertheless offers a strong comparative compliance benchmark for designing recording and human intervention together.
When a student appeals an artificial intelligence-supported grade, the easiest technical solution is to send the answer to the same model again and obtain a second score. This may be useful as a consistency test; but it is not re-examination on its own. A second output tied to the same error pattern, the same bias, or the same wrong rubric interpretation cannot count as independent control of the first decision.
An effective appeal process should have eight stages:
Accessible application: The student and, where appropriate, the parent should be able to apply through a short, understandable, age-appropriate channel without excessive formal requirements.
Preservation of the file: Original work, rubric, first model output, and human action records must be kept in an unalterable form.
Scope of application: The student should not be forced to say only “the model made a mistake”; they should also be able to raise academic assessment, the criterion used, accuracy of data, discriminatory effect, accessibility problem, or procedural deficiency.
Independent first look: Especially in high-impact decisions a second assessor should, where possible, score the original answer against the rubric before seeing the artificial intelligence score. This reduces the anchoring effect of the first number.
Comparison and investigation: Independent assessment should be compared with the first outcome; if there is a significant difference, rubric, model output, contextual information, and records should be examined.
Genuine power to change: The appeal authority must be able to change the grade, reverse the action, or order new assessment. Vendor support may be obtained; final review cannot be left to the vendor.
Justified outcome: The student should be told how each claim was assessed, whether the outcome changed, and if applicable the next application route.
Propagation of correction: The change must not remain only in the grade book; success profile, scholarship, level, risk indicator, passing, or notification systems fed by this grade must also be corrected.
MEB measurement and assessment documents offer an important domestic basis for this architecture. Written and practical examinations provide for showing examination papers to the student, a time-limited appeal route on results, and feedback. In addition, open-ended questions in common written examinations conducted electronically or by scanning answer sheets into digital form are assessed by two independent scorers; where disagreement cannot be resolved, recourse to a senior scorer is provided.[10]
This double-scoring rule is not a general legal rule automatically applicable in every teacher examination or every use of artificial intelligence. But it makes visible an important procedural idea: In open-ended educational assessment, reliability can be produced not only from the first assessor’s title but from comparing independent views and a dispute-resolution mechanism. The same idea can be transferred to institutional policy in proportion to risk for high-impact artificial intelligence-supported decisions.
Independent review does not mean establishing a separate panel for every short exercise. The intensity of safeguards should increase according to the decision’s impact on the student, reversibility, artificial intelligence’s role, and the likelihood of error spreading.
Low-impact decisions are uses such as formative feedback, practice question, or suggestion; advance information on artificial intelligence use, easy correction by the teacher, error reporting channel, and sample-based quality control suffice. Grade-affecting decisions contribute to homework, project, in-class examination, or term grade; per-student human review, rubric and brief justification, record of first decision, timely appeal, and where possible a different assessor are required. High-impact decisions cover passing a class, graduation, programme admission, scholarship, discipline, or significant level placement; stronger independence, panel or senior assessor if needed, detailed decision trail, swift interim measure, justified decision, and correction of all derivative outcomes are sought.
The second assessor’s independence should not be measured only by the organisation chart. A person whose performance is evaluated for defending the first decision, who is subject to the same bulk-approval targets, or who sees only output rerun by the vendor may not be functionally independent. In a small institution review of certain files by another teacher, department head, or panel on a risk basis may suffice. The criterion is a genuine second look authorised to reassess, not to confirm the first outcome.
In high-impact decisions with approaching deadlines, waiting for the appeal outcome itself may cause harm that is hard to remedy. If scholarship application period, registration renewal, graduation, or disciplinary sanction is at stake, the institution should define in its procedure measures such as temporarily suspending the effect of the decision, conditional treatment of the student, or expedited review. Effective application is not one that exists only in theory but one that works while the outcome can still be corrected.
The student cannot always be expected to diagnose the source of technical error. The application form should allow raising one or more of the following claims:
application of a wrong or different version of the answer key or rubric;
failure to recognise the correct answer because of language, form of expression, or alternative solution method;
the model taking into account data not belonging to the answer or producing an element not present in the answer;
failure to reflect disability, special education need, linguistic difference, or accommodation in assessment;
personal data being wrong, outdated, or irrelevant;
use of an unauthorised tool or unapproved version;
no human review or review limited only to bulk approval;
inexplicably different scores for answers of similar quality;
failure to give prior notice of artificial intelligence use or insufficient justification;
another deficiency relating to conflict of interest, impartiality, or procedural rules.
This list can also become the institution’s incident classification. Closing appeals only as “justified/unjustified” destroys information valuable for compliance. If it is seen which model version, question type, and student groups concentrate which types of errors, use may be narrowed, the rubric changed, teacher training renewed, or the tool stopped entirely.
In a dispute over an artificial intelligence-supported grade, academic appeal, YAZEK ethical notification, KVKK application, and administrative or contractual routes must not be conflated. Their purpose, addressee, and outcome differ.
Grade or examination appeal aims at re-examination of academic assessment and procedure; possible outcomes are rescoring, retention or change of grade, or new assessment. YAZEK ethical notification process examines compliance of artificial intelligence use with MEB’s ethical framework; ethical review, corrective measure, and referral to a higher board may follow, but this route cannot be assumed to change the grade on its own. KVKK data-subject application is for information on personal data processing, access, correction, and where conditions exist objection to automated outcome; the data controller’s response, data correction/deletion, or reassessment and application and complaint routes under the Law come into play. Administrative, judicial, or contractual route reviews compliance of the decision with law, institutional regulations, and the relationship between the parties; depending on the concrete regime, cancellation, correction, compensation, or other outcomes may follow.
The ethical notification and evaluation chain in YAZEK is important for addressing ethical appropriateness of artificial intelligence use within the institution.[11] But this chain does not automatically suspend a grade appeal period and does not replace academic reassessment. Likewise a KVKK application is not in every case the authority that decides the academic correctness of the grade. The institution may offer the student a “single door”; but it must direct the application internally to the right processes and must not create the impression that one route prevents resort to another.
Even if KVKK Art. 11/1-g conditions are not met, other rights such as learning whether one’s data are processed, requesting information, learning whether use is consistent with purpose, and requesting correction of wrong data may remain in play.[12] Examination legislation, internal regulations, student contracts, and public-law safeguards must also be assessed separately according to the concrete institution. Application design should therefore not be tied only to the “is there an automated decision?” threshold.
A well-designed appeal mechanism does not only resolve individual disputes; it shows the real-world performance of the institution’s artificial intelligence use. Laboratory tests show what the model does on expected examples; appeals tell what the student experienced in unexpected situations.
Without unnecessarily duplicating personal data or creating new risks for groups, the institution can monitor:
number of appeals and ratio by decision type;
rate of overturned or corrected decisions;
magnitude of correction on the score;
difference between first and second assessor;
number of decisions without justification or with incomplete record;
error clusters by model, version, question type, and course;
distribution of accessibility, linguistic difference, or discrimination claims;
time to conclusion of application;
whether correction reached derivative systems.
These indicators must not be turned into a new performance system that automatically scores teachers or students. Such secondary use additionally requires purpose, legal basis, proportionality, and fairness assessment. The aim here is not to rank individuals but to catch recurring risks in the process.
If overturn rate rises in a particular model version, use may be suspended. If scorers consistently diverge on certain question types, the rubric may be reviewed. If humans almost never deviate from the model suggestion, the reason may be interface design or workload rather than extraordinary accuracy. Appeal data therefore bridges compliance monitoring, model change management, and the currency of the institution’s YAZEK declaration.
If human oversight, justification, and appeal are scattered across different policies, gaps arise in practice. A separate Artificial Intelligence-Supported Educational Decision and Appeal Procedure linking the educational institution’s information security, personal data, measurement-assessment, and artificial intelligence policies can close this gap.
This procedure should cover at least the following eight headings:
Scope and decision inventory: In which course, examination, assignment, discipline, admission, or guidance decision which artificial intelligence function is used.
Impact classification: Importance of the decision, reversibility, student age, possibility of special-category data, and de facto determinativeness of the model.
Authority matrix: Who may use the system, who gives the final decision, who reviews appeal, and who may stop the system.
Human review standard: Access to actual answer, rubric use, training, time, power to change, and intervention record.
Explanation standard: Separate contents of advance disclosure, student-specific justification, and limited oversight file.
Decision trail and retention: Which records, for how long, with which security measures and access rights.
Appeal and correction: Application channel, period, independence, interim measure, decision format, and propagation of correction to all linked systems.
Monitoring and change management: Appeal indicators, model/version changes, retesting, incident management, narrowing or stopping use.
The vendor contract should also support this procedure. A product to which the institution lacks access to model and version information, actual output records, change notification, and technical support needed for review may make effective appeal impossible even if it offers “human approval” on screen. The contract alone is nevertheless insufficient: User screen, teacher workload, examination rules, and internal delegation of authority must be aligned to the same standard.
MEB’s ethical guide offers a framework for education, teaching, and management processes in the Ministry’s central, provincial, and overseas organisation and in official and private pre-school, primary, and secondary institutions.[13] Universities should not be assessed within YAZEK’s this institutional scope. In higher education, decision authority, examination and appeal regulations, public or foundation university character, and personal data processes must be examined separately through their own legal documents.
The basic architecture built in this article nevertheless has broader value. If artificial intelligence produces an output affecting a student’s grade, discipline, scholarship, access to programme, or academic progress, the institution must be able to answer four questions:
How effective was the system in the decision in reality?
What original assessment did the authorised human make?
Could the student understand the reason for their own outcome?
Could error be corrected effectively and independently of the system that produced the outcome?
These questions may take different forms in private school through contract and consumer relations, in public school through education legislation and administrative operation, in university through internal academic regulations. But none should remain merely an “ethical principle”; each must be converted into a procedure that works through duty, record, period, and authority.
In artificial intelligence-supported assessment, reliability does not arise only from the model’s average accuracy rate. For the student, the reliable system is one that can show by which criterion their own answer was assessed, does not hide the model’s role, proves that the authorised human actually decided, and offers effective second look when there is error.
For this reason the answer to “Who gave the grade?” cannot consist of a username. The right answer must be the authorised human who saw the actual work, could independently evaluate the artificial intelligence suggestion, had authority to change the outcome, built student-specific justification, and can account for this decision. The institution must not hide behind this human; it must provide the necessary time, training, interface, record system, and appeal authority.
A grade is not the number the system produced. A grade is an institutional outcome that can explain which learning outcome was assessed on which evidence in which student answer. The same model’s answer a second time is not second review either. Real safeguard is a human eye that can think outside the technology and change the first outcome.
For YAZEK compliance, institutions’ next step is to test their general ethical declarations with a decision and appeal gap analysis: At every point where artificial intelligence touches the grade, de facto determinativeness, human ownership, student-specific justification, decision trail, and propagation of correction must be tested. The Artificial Intelligence-Supported Educational Decision and Appeal Procedure prepared from this analysis can unite the educational institution’s artificial intelligence policy, measurement-assessment arrangement, KVKK processes, and vendor controls on the same practicable map.
Editorial note: This work is prepared for general information purposes. It does not replace legal opinion for a particular institution or dispute. Applicable legislation and application route must be assessed separately according to the institution’s character, type of decision, function of the system used, and concrete data processing activity.