diff --git a/RV-Monitor/Kernel_Safety_Claims_by_Runtime_Verification_Monitors.md b/RV-Monitor/Kernel_Safety_Claims_by_Runtime_Verification_Monitors.md new file mode 100644 index 0000000..e597a67 --- /dev/null +++ b/RV-Monitor/Kernel_Safety_Claims_by_Runtime_Verification_Monitors.md @@ -0,0 +1,542 @@ + + +# Kernel Safety Claims by Runtime Verification Monitors + + +# Introduction + +The goal of this document is to define a comprehensive methodology to design an RV Monitor against a set of allocated safety requirements. \ +Instead we do not intend to prove the RV Monitor to be able to accommodate every single use case. + + +### Why it is hard to qualify the Kernel to support associated safety claims + +The Linux Kernel is monolithic and provides thousands of functionalities in order to support a huge amount of applications across multiple industry domains and across many different target HW platforms. + +For this reason the Kernel is designed to be modular (see [MAINTAINERS](https://github.com/torvalds/linux/blob/master/MAINTAINERS)), however each Kernel element (subsystem or driver) provides many functionalities and all such elements can interact with each other in a very complex way. To understand how complex such interactions can be, fig. [1] in the Appendix shows the possible call tree in the context of ioctl between “FILESYSTEMS (VFS and infrastructure)” and the drivers or subsystems it communicates with. \ + \ +As an example, if there is a safety requirement associated with ioctl() (e.g. ioctl shall be used to set the correct timeout in an external safety watchdog device), even if the operation is in principle very simple, the Kernel code providing the implementation must be qualified to make sure: + + + +1. The kernel code implementing the requirement is correct; +2. There is no unintended functionality in the current implementation; +3. There is no undetected interference from other Kernel code outside the current implementation. + +With respect to points 1. and 2. above we can classify Kernel code into two classes: + + + +1. Kernel code that functionally contributes to the allocated safety requirement +2. Kernel code potentially invoked as part of the implementation but not contributing to the safety requirement (e.g. printk). + +With respect to point 3. above we can have: + + + +3. Unintended functionalities due to interferences from Kernel code outside the implementation associated with classes A. and B. above. + + +When it comes to claim that code as in A. and B. are able to meet point 1 and point 2, for a complex SW component we need to hierarchically break it down into SW units, for each of these do a safety analysis and derive CoUs and Derived Safety Requirements; this regardless of code being classified as A. or B above. This process already proved to be very intensive even considering an entire subsystem like a single SW Unit (e.g. [STPA(-like) inside the Kernel ](https://docs.google.com/document/u/0/d/1K_cQSS2KYDnJQ0B91Zvlxq9-35Cx8ntbXwVMQJ51rvY/edit)). + +Finally with respect to failure modes as in C it is well known that, as of today, there are no comprehensive architectural measures that can isolate Kernel implementing the allocated safety requirement from code outside the implementation associated with A. and B. + + +### Which problems are solved by runtime verification monitors + +[Runtime Verification Monitors](https://docs.kernel.org/trace/rv/index.html) are mechanisms inside the Kernel that allow monitoring the Kernel itself to behave according to a predefined machine state diagram. \ +So if the Kernel misbehaves leading to a violation of points 1 or 2 above, the runtime verification monitor detects that the monitored Kernel code has violated the associated machine state diagram and accordingly raises an exception halting the execution or doing other custom remedial actions. \ +In this regard RV Monitors could be used not only to do control flow monitoring but also data integrity monitoring. In fact it may be possible to associate states with corresponding valid data sets that could be stored internally as part of the monitor specific data structures and, upon incoming events, it may be possible to check the correctness of the event specific data against the current state and associated data set. + +So assuming that it is possible to prove FFI between RV Monitors and the rest of the Kernel code (see section below), adopting the RV Monitor solution would lead to the following advantages: + + + +* The hierarchical safety analysis of the Kernel and associated safety requirements definitions can be limited to code functionally contributing to the allocated safety reqs (class A above). +* The Runtime Verification Monitor provides an architectural protection mechanisms against Kernel interferences following a safety analysis of the resources to be monitored against interference failure modes + + +### When is it valuable to use RV Monitors? + +From a conceptual perspective the RV Monitors introduce redundancy inside the Kernel, it is therefore valuable to use them as qualification measure when the following criteria are met: + + + +* The complexity of the Automata model is significantly lower than the code being monitored (i.e. when the code is capable of providing many functionalities under many different condition and states but only a specific subset of its potentials are used in a safety context); this implies that from a FuSa qualification effort perspective and continuous certification perspective dealing with a significantly smaller code baseline is cheaper and faster and hence it is better to make a systematic capability claim on the RV Monitor code instead of spending effort on proving the same for the monitored code. +* The expected behavior of the monitored code is comprehensively analyzed in how it would meet the allocated safety requirements. This is needed to guarantee that the monitor is monitoring all the critical states and data for safety reasons and to avoid the monitor triggering due to valid states and data not being analyzed. +* The specific use case being evaluated is compatible with ex-post-facto treatment of the failure: by design the RVM will be able to detect the failure only after it has taken place and it cannot avoid it (it works as detection measure, not as prevention one) . +* It is possible to determine that the RVM shall always trigger before the detected failure propagates in a dangerous way to the safety workloads being monitored (avoid failure propagation) + + +# Integration of Runtime Verification Monitors with the Kernel + +A monitor is the central part of the runtime verification of a system. The monitor stands in between the formal specification of the desired (or undesired) behavior, and the trace of the actual system. + +In Linux terms, the runtime verification monitors are encapsulated inside the RV monitor abstraction. A RV monitor includes a reference model of the system, a set of instances of the monitor defined at development time (e.g. per-cpu monitor, per-task monitor, and so on), and the helper functions that glue the monitor to the system. Depending on both the parameters being monitored and the specific hardware features available, different approaches can provide flexibility in managing the need of monitoring vs the overhead that it might introduce. More on this later. + +Callbacks must be introduced in a way that is compatible with safety requirements allocated to the functionality they support (i.e. adding tracepoint by patching the code rather than by compiler could be more cost-effective from a FuSa effort point of view). + +Currently, the RV subsystem allows the usage of automata as formalism. The formalism uses **states** and **events** as entry points. Each time a callback is called, it produces an **event**, that is checked in the **current** **state**. The processing of the **event** and **state** results in the **next state**,** **which will be saved. + +The state transition table used to define how the **current state** changes according to the possible **events** is a read only **matrix**. The current state is saved in a writable memory accessible in **kernel space**. + +In addition to the verification and monitoring of the system, a monitor can react to an unexpected event. If an **event** is not valid for the **current state**, the automata will report an **unexpected event**. The forms of reaction can vary from logging the event occurrence to the enforcement of the correct behavior to the extreme action of taking a system down to avoid the propagation of a failure. + +The [current state ](https://git.kernel.org/pub/scm/linux/kernel/git/bristot/linux.git/tree/include/linux/rv.h#n18)variable is stored in a regular memory space in the kernel. The place varies according to the **monitory type**. It can be either in a **global variable** for the global monitor, on a **per-cpu variable** on a per-cpu monitor, or data stored in the **task struct** for per task monitor. + +The same consideration is true for the [monitoring](https://git.kernel.org/pub/scm/linux/kernel/git/bristot/linux.git/tree/include/linux/rv.h#n17) variable; such variable is used to activate or deactivate a specific instance of runtime verification monitor. + +On top of such variables (that are defined as part of the da_monitor structure) it is possible to define any custom monitor specific data set that can be associated with the monitor states and be checked against event specific data to make sure the event itself is valid and the data set coming with it is also valid when compared against the current state specific data set. + + +# Freedom From Interference between RV Monitors and the rest of Kernel code + +From an interference perspective we have different types of interference: + + + +* **Temporal**: Kernel code slowing down the monitor code so much to miss the associated FDTI deadline +* **Communication**: Kernel code interfering with the monitor code through the interfaces between the Kernel and the RV monitors +* **Spatial**: Kernel code interfering with the monitor code through random corruption of the Kernel address space + +From a **temporal interference** point of view we can assume the safety relevant code (including the Kernel code) to be monitored by an external watchdog. So if such code slows down beyond the allocated max FDTI portion, the external watchdog will trigger. \ +Now the monitored code by definition is safety relevant code, and, as explained above, the associated RV monitor executes in the same context directly invoked by the trace interfaces added in the monitored code. \ +So if the monitored code slows down, the RV Monitor code equally slows down and as a result the external watchdog is not pet on time, with a safe state driven physically by the watchdog. + +Even in case of failure being detected by the RV Monitor, since the RV Monitor code executes synchronously with respect to the monitored code, if a fault reaction is not completed on time, in the end the external watchdog will trigger before the allocated max FDTI portion expires. + +From a **communication** interference point of view, the only interfaces between the Kernel code and the RV monitors are the trace interfaces added by design in the code being monitored. So by design the lack of communication interference is guaranteed by the process that is followed to actually design the instances of RV Monitors (as further elaborated below). + +More specifically, failures in the code being monitored leading to wrong data being passed to the trace interfaces should already be covered by the safety analysis that precedes the design of the specific RV Monitor instance covering such code. + +From a **spatial interference** point of view there are different failures modes that we need to consider: + + + +* **FM1:** the event or event specific data passed to the runtime verification monitor instance is corrupted by a spurious stack corruption \ +**Effect of Failure: **the event corrupted could lead the RV monitor in a valid state that is not however the one of the code being monitored. \ +**Prevention or detection measures:** The RV Monitor shall be designed to detect corruption of relevant DATA due to bugs in the code being monitored. \ +Therefore there would be no difference, from a FuSa perspective, between a corruption of parameters due to a bug in the monitored Kernel code and a corruption of the same parameters due to random stack access from QM or NSR Kernel code. \ +So by design the RV Monitor shall be able to detect such failure mode. \ + \ +**Note**: a wrong parameter in the monitored code that turns to a valid one due to corruption on the stack is considered a double point of failure, hence not needed/ in scope. + +* **FM2:** the current state or state specific data get corrupted on any writable memory region mapped in the Kernel address space due to random interference from Kernel code. \ +**Effect of Failure: **the current state could switch to a wrong valid state or the state specific data could be dangerously change in a way that could lead the monitor to be ineffective. \ +**Prevention or detection measures:** the current state and state specific data can be CRC protected before storing them, hence any random corruption would lead to a CRC check failure as the Monitor code is triggered by the next event (CRC could be replaced by less expensive integrity check algorithms if the targeted ASIL allows). +* **FM3: **RV Monitor code is corrupted due to random interference from Kernel code. \ +**Effect of Failure: **RV Monitor code could not execute or wrongly execute leading to the monitor being ineffective. \ +**Prevention or detection measures: **RV Monitor code is stored in memory mapped with the RO attribute, hence any spurious write would trigger an SMMU exception that would, at least, halt the execution of the Kernel, hence resulting in the external watchdog to trigger and drive the platform safe state + + +# RV Monitors Design + + +### <from failure modes to automata finite state diagrams> + + +### + + +### https://docs.kernel.org/trace/rv/deterministic_automata.html + + +### <from automata to RV monitor implementation> + + +# RV Monitors Examples + + +## i6300esb Watchdog timeout monitor + + +### Introduction and Requirements + +The goal of this RV Monitor is to qualify the Kernel according to a subset of the safety requirements defined for the Telltale Safety Application in this [sheet](https://docs.google.com/spreadsheets/d/1EbuVvhXo-xZc2aPTfMgQtPNPDQYtcozs/edit#gid=584539121). More specifically: + + + +* **KSR_0003**: the watchdog subsystem shall ensure the opened WD device to be the one specified as input argument +* **KSR_0004**: The watchdog subsystem shall ensure the WD timeout to be set according to the IOCTL input parameter +* **KSR_0005**: The operating system shall ensure the WD timeout to be not wrongly re-set to a different timeout value +* **KSR_0012**: Writing the Watchdog shall reset it with a timeout equal to the one specified in the IOCTL. No other timeout value shall be used / reconfigured + +**Assumption:** it is assumed that any boot time operation has been successfully completed; we could consider the design of a RV Monitor for guarding against boot time misconfigurations as a follow-on activity + + +### Static and Dynamic Architecture + + + \ +As a first step we consider only the architectural elements that are functionally supporting the implementation of the safety requirements mentioned above; interfering architectural elements will be added later in the picture. \ + \ +In the use case under analysis, to start with, we can identify 3 main elements interacting with each other: + + + +* The Safety Application +* The Linux Kernel +* The HW SoC that also includes a programmable i6300esb watchdog. + + \ +Accordingly we have the following events that characterize the expected interactions in absence of failures. + + + +1. The Safety Application invokes the open syscall passing the input path from the filesystem associated with the watchdog to be opened. +2. Based on the path passed from open, the Kernel: + 1. Allocates a unique file descriptor (fd) associated with the such watchdog device; + 2. Unlocks the watchdog registers for write access; + 3. Reloads the external watchdog (stops any previous counting) + 4. Starts the watchdog using the default timeout programmed at boot time; +3. In case of success the Kernel returns the allocated unique fd to the Safety Application +4. The safety application uses the returned fd to set a new watchdog timeout by invoking `ioctl(fd, WDIOC_SETTIMEOUT, &timeout`); where: + 5. fd is the watchdog file descriptor returned by open + 6. `WDIOC_SETTIMEOUT `is the ioctl command that corresponds to the operation of setting a watchdog timeout + 7. `&timeout `is the user space address of the variable containing the timeout value to be written to the watchdog +5. According to the fd passed by the Safety Application the Kernel: + 8. Shift Left by 9 bits the timeout value; + 9. Unlocks the watchdog registers for write access; + 10. Write the shifted value to `ESB_TIMER1_REG;` + 11. Unlocks the watchdog registers for write access; + 12. Write the same shifted value to `ESB_TIMER2_REG;` + 13. Unlocks the watchdog registers for write access; + 14. Reloads the external watchdog (stops any previous counting) +6. If successful the Kernels returns 0 to the Safety Application, else -1 is returned. +7. The Telltale apps starts the safety operations; the watchdog is pet by writing to the i6300esb watchdog device +8. The Kernel unlocks the watchdog registers for write access and reloads the watchdog timer +9. The write operations returns a non negative number to the safety application + + + +### Definition of the i6300esb RV Monitor + +At this stage there are two possibilities: + + + +1. A safety analysis is done on the Kernel behavior associated with the architecture described above and an initial RV Monitor design is conceptualized. Then both design and safety analysis are updated till all dangerous failure modes are under control. +2. An initial RV monitor is designed and a safety analysis is done to verify its effectiveness against failure modes of the Kernel. Then both design and safety analysis are updated till all dangerous failure modes are under control. + +In the end from an end result point of view there is no difference between them. With respect to this specific example we’ll start with an initial RV Monitor design that will be later verified by safety analysis and possibly refined. + + +#### i6300esb RV Monitor Initial Safety Requirements + + + +* **KSR_RVM_0003**: following a call to open() with the input path associated with the i6300 watchdog device, the RV Monitor shall trigger kernel panic if a valid fd is returned by open and the driver either misses to do any expected operation on the i6300esb HW or does it in the wrong order +* **KSR_RVM_0004**: following a call to ioctl with the same fd returned by open and with `WDIOC_SETTIMEOUT` as input cmd, the RV Monitor shall trigger panic if success is returned and the driver either misses to do any expected operation on the i6300esb HW or does it in the wrong order; +* **KSR_RVM_0005**: after a successful invocation of ioctl() the RV Monitor shall trigger Kernel Panic if the i6300esb HW registers to set the timeout value (ESB_TIMER1_REG, ESB_TIMER2_REG) are attempted to be written +* **KSR_RVM_0012: ** after a successful invocation of ioctl() the RV Monitor shall trigger Kernel Panic if any thread different from the one associated with the Telltale Application unlock the write access to the watchdog registers +* **KSR_RVM_0012.1: ** after a successful invocation of ioctl() the RV Monitor shall trigger Kernel Panic if the writes invoked from the Telltale Application returns success and the driver either misses to do any expected operation on the i6300esb HW or does it in the wrong order; +* **KSR_RVM_0012.2: ** after a successful invocation of ioctl() the RV Monitor shall trigger Kernel Panic if any write is attempted to ESB_LOCK_REG + + +#### i6300esb RV Monitor Initial Design + + + + + + +The RV Monitor is modeled by the automata diagram pictured above. Each circle represents a state of the RV monitor (monitoring the whole Kernel in this case), each arrow represents an event that leads to a state transition. + + + +
| State + | +State Description + | +Possible Events + | +Next State + | +
| init + | +This is the initial state following the system boot up + | +open: the open syscall has been invoked passing the FS path associated with the i6300 watchdog device + | +start + | +
| start + | +open has been invoked and accordingly the i6300 watchdog is enabled and starts ticking according to the default timeout + | +failure: something went wrong enabling or starting the i6300 watchdog + | +init + | +
| success: Open is going to return a valid file descriptor + | +start_check + | +||
| start_check + | +Before open returns success, the RV monitor checks if the operations on the HW have been performed as expected + | +success: the operations on the HW have been performed as expected + | +started + | +
| failure: open was going to return success while one or more operations on the HW were not as expected + | +Fail + | +||
| started + | +the i6300 watchdog is enabled and is ticking according to the default timeout, the corresponding file descriptor has also been returned to the safety application + | +ioctl: the ioctl syscall has been invoked on the i6300 file descriptor with the command WDIOC_SETTIMEOUT and a timeout value
+ |
+ heartbeat + | +
| heartbeat + | +ioctl has been invoked and the i6300 watchdog timer registers are programmed according to the timeout value passed by the safety app. Also the watchdog starts ticking + | +failure: something went wrong with the ioctl operation + | +started + | +
| success: the ioctl operation succeeded and ioctl is going to return success to the safety app + | +heartbeat_check + | +||
| heartbeat_check + | +Before ioctl returns success, the RV monitor checks if the operations on the HW have been performed as expected + | +success: the operations on the HW have been performed as expected + | +beating + | +
| failure: ioctl was going to return success while one or more operations on the HW were not as expected + | +Fail + | +||
| BPA / unlock_wr: the watchdog timer registers were attempted to be written or a thread different from the safety app tried to unlock the i6300 HW for write access + | +Fail + | +||
| beating + | +The ioctl returned success and the i6300 watchdog is ticking according to the programmed timeout + | +write: the Safety App invoked a write syscall on the i6300 file descriptor + | +reload + | +
| BPA / unlock_wr: the watchdog timer registers were attempted to be written or a thread different from the safety app tried to unlock the i6300 HW for write access + | +Fail + | +||
| reload + | +The watchdog is reloaded according to the current timeout and starts ticking + | +failure: something went wrong when reloading the watchdog + | +beating + | +
| success: the reload operation succeeds and write is going to return success + | +reload_check + | +||
| BPA / unlock_wr: the watchdog timer registers were attempted to be written or a thread different from the safety app tried to unlock the i6300 HW for write access + | +Fail + | +||
| reload_check + | +Before write returns success, the RV monitor checks if the operations on the HW have been performed as expected + | +success: the operations on the HW have been performed as expected + | +beating + | +
| failure: write was going to return success while one or more operations on the HW were not as expected + | +Fail + | +||
| BPA / unlock_wr: the watchdog timer registers were attempted to be written or a thread different from the safety app tried to unlock the i6300 HW for write access + | +Fail + | +||
| fail + | +The monitored code is in an invalid state, hence the RV Monitor triggers Kernel Panic here + | +None + | ++ | +