International Conference on Rebooting Computing (ICRC) 2016
ICRC 2016 sought to discover and foster novel methodologies to reinvent computing technology, including new materials and physics, devices and circuits, system and network architectures, and algorithms and software.
On this Event Showcase page, you will see some of the featured presentations from the researchers who are leading this growing conference and community.
Molecular Cellular Networks: A Non von Neumann Architecture for Molecular Electronics - Craig Lent: 2016 International Conference on Rebooting Computing
The two fundamental limitations of the present computing paradigm are power dissipation from transistor switching and the architectural von Neumann bottleneck that segregates processing from memory. We examine a cellular architecture which radically intermixes memory and processing, and which is based on a transistor-less approach to representing binary information using the arrangement of charge within the molecule. Representing bits by molecular configuration, rather than a current switch, yields the limits of functional density and low power dissipation. Matching a new computational element to a new architectural framework could enable general purpose computing to evolve along a new roadmap.
Computing with Dynamical Systems - Fred Rothganger: 2016 International Conference on Rebooting Computing
The effort to develop larger-scale computing systems introduces a set of related challenges: Large machines are more difficult to synchronize. The sheer quantity of hardware introduces more opportunities for errors. New approaches to hardware, such as low-energy or neuromorphic devices are not directly programmable by traditional methods. These three challenges may be addressed, at least for a subset of interesting problems, by a dynamical systems approach. The initial state of system represents the problem, and the final state of the system represents the solution. By carefully controlling the attractive basin of the system, we can move it between these two points while tolerating errors, which appear as perturbations. Here we describe both conventional and neural computers as dynamical systems, and show how to construct algorithms with resilience to noise, using traditional numerical problems as a special case. This suggests a reduction from numerical problems to spiking neural hardware such as IBM's TrueNorth.
Towards Logic-in-Memory circuits using 3D-integrated Nanomagnetic Logic - Fabrizio Riente: 2016 International Conference on Rebooting Computing
Perpendicular NanoMagnetic logic (pNML) is one emerging beyond-CMOS technology listed in the ITRS roadmap for next-generation computing due to its non-volatility, monolithic 3D-Integration, small size and low power consumption. Here, we demonstrate the feasibility of a monolithic 3D pNML circuit, which is capable of integrating both memory and logic onto the same device on different layers exploiting the novel Logic-In-Memory (LIM) concept. The LIM can be exploited by placing magnetic memory elements (registers) in a memory layer, which is located monolithically just below the performing logic plane and interconnected by pure-magnetic vias. In particular, the nonvolatile magnetization state of the bistable, nanoscaled magnets with perpendicular magnetic anisotropy is exploited to build a magnetic D flip-flop. This basic memory element is then used to build a more compact and a more power efficient N-bits parallel-in parallel-out registers. Indeed, the presented magnetic flip-flop implementation is two orders of magnitude more compact when compared to the 32nm CMOS version. The approach has been studied by considering the implementation of an accumulator (adder plus memory) as case study. This novel concept allows the storage of information locally on the computing chip, saving area and employing the strengths of pNML for next-generation, memory-intensive computing tasks.
Conversion of Artificial Recurrent Neural Networks to Spiking Neural Networks for Low-power Neuromorphic Hardware - Emre Neftci: 2016 International Conference on Rebooting Computing
In recent years the field of neuromorphic computing gained significant momentum, enabling systems that consume orders of magnitude less power than traditional ones. However, their wider use is still hindered by the lack of algorithms that can harness their full potential. Recurrent neural networks (RNN) are widely used in machine learning to solve a variety of sequence learning tasks. In this work we present a "train-and-constrain" methodology that enables the mapping of machine learned RNNs to spiking neurons. This "train-and-constrain" method consists of first training RNNs, then discretizing the weights and finally converting them to spiking RNNs. We demonstrate our approach by mapping a natural language processing task (question classification), where we demonstrate the entire mapping process of the recurrent layer of the network on IBM's Neurosynaptic System TrueNorth, a spike-based digital neuromorphic hardware architecture (including adapting the network to constraints associated with the system). Surprisingly, we find that short synaptic delays are sufficient to implement the dynamic (temporal) aspect of the RNN in the question classification task. The hardware-constrained model achieved 74% accuracy in question classification while using less than 0.025% of the cores on one TrueNorth chip, resulting in an estimated power consumption of ~17 uW.
Spiking Network Algorithms for Scientific Computing - William Severa: 2016 International Conference on Rebooting Computing
For decades, neural networks have shown promise for next-generation computing, and recent breakthroughs in machine learning techniques, such as deep neural networks, have provided state-of-the-art solutions for inference problems. However, these networks require thousands of training processes and are poorly suited for the precise computations required in scientific or similar arenas. The emergence of dedicated spiking neuromorphic hardware creates a powerful computational paradigm which can be leveraged towards these exact scientific or otherwise objective computing tasks. We forego any learning process and instead construct the network graph by hand. In turn, the networks produce guaranteed success often with easily computable complexity. We demonstrate a number of algorithms exemplifying concepts central to spiking networks including spike timing and synaptic delay. We also discuss the application of cross-correlation particle image velocimetry and provide two spiking algorithms; one uses time-division multiplexing, and the other runs in constant time.
Accelerating Machine Learning with Non-Volatile Memory: Exploring device and circuit tradeoffs - Pritish Narayanan: 2016 International Conference on Rebooting Computing
Large arrays of the same nonvolatile memories (NVM) being developed for Storage-Class Memory (SCM) -- such as Phase Change Memory (PCM) and Resistance RAM (ReRAM) -- can also be used in non-Von Neumann neuromorphic computational schemes, with device conductance serving as synaptic "weight." This allows the all-important multiply-accumulate operation within these algorithms to be performed efficiently at the weight data. In contrast to other groups working on Spike-Timing Dependent Plasticity (STDP), we have been exploring the use of NVM and other inherently-analog devices for Artificial Neural Networks (ANN) trained with the backpropagation algorithm. We recently showed a large-scale (165,000 two-PCM synapses) hardware-software demo (IEDM 2014) and analyzed the potential speed and power advantages over GPU-based training (IEDM 2015). In this paper, we extend this work in several useful directions. We assess the impact of undesired, time-varying conductance change, including drift in PCM and leakage of analog CMOS capacitors. We investigate the use of non-filamentary, bidirectional ReRAM devices based on PrCaMnO, with an eye to developing material variants that provide suitably linear conductance change. And finally, we explore tradeoffs in designing peripheral circuitry, balancing simplicity and area-efficiency against the impact on ANN performance.
Opportunities in Physical Computing driven by Analog Realization - Jennifer Hasler: 2016 International Conference on Rebooting Computing
In the past, discussions on the capability of analog or physical computing were only of theoretical interest. Digital computation's 80 year history starts from the Turing's original model of computation to ubiquitous modern computational devices. The modern development of analog computation started with almost zero computational framework. Today, we have significant programmable and configurable physical computing systems. The focus of this paper is to have these discussions given the very real potential of ultra-low power physical computing systems. This work considers the current state of analog computation, energy efficient computation, and analog numerical analysis, moving towards starting a unified analog-computing framework, including quantum computing, as part of physical computing.
Neuromorphic Mixed-Signal Circuitry for Asynchronous Pulse Processing Neuromorphic Mixed-Signal Circuitry for Asynchronous Pulse Processing - Peter Petre: 2016 International Conference on Rebooting Computing
We demonstrate a software reconfigurable mixed-signal Printed Circuit Board (PCB) prototype and a custom mixed-signal Application Specific Integrated Circuit (ASIC) prototype of a cognitive signal processor using neuromorphic methods to perform adaptive nonlinear filtering based real-time wideband signal processing algorithms. The cognitive processor effectively implements a trending computing paradigm called Reservoir Computer (RC). Hardware implementation of the RC is achieved by a novel analog signal processor architecture called the Asynchronous Pulse Processor (APP).
Accelerating Discrete Fourier Transforms with Dot-product engine - Miao Hu: 2016 International Conference on Rebooting Computing
Discrete Fourier Transforms (DFT) are extremely useful in signal processing. Usually they are computed with the Fast Fourier Transform (FFT) method as it reduces the computing complexity from O(N^2) to O(Nlog(N)). However, FFT is still not powerful enough for many real-time tasks which have stringent requirements on throughput, energy efficiency and cost, such as Internet of Things (IoT). In this paper, we present a solution of computing DFT using the dot-product engine (DPE) a one transistor one memristor (1T1M) crossbar array with hybrid peripheral circuit support. With this solution, the computing complexity is further reduced to a constant O() independent of the input data size, where is the timing ratio of one DPE operation comparing to one real multiplication operation in digital systems.
Digital Neuromorphic Design of a Liquid State Machine for Real-Time Processing - Nicholas Soures: 2016 International Conference on Rebooting Computing
The Liquid State Machine (LSM) is a form of reservoir computing which emulates the brains capability of processing spatio-temporal data. This type of network generates highly descriptive responses to continuous input streams. The response is then used to extract information about the input stream. A single LSM network can be used as a generic intelligent processor that processes different streams of data (or) on same stream of data to extract different features. The LSM has been shown to perform well in tasks dependent on a systems behavior through time. The LSM's intrinsic memory and its reduced training complexity make it a suitable choice for hardware implementations for spatio-temporal applications. Existing behavioral models of LSM cannot process real time data due to their hardware complexity or inability to deal with real-time data or both. The proposed model focuses on a simple liquid design that exploits spatial locality and is capable of processing real time data. The model is evaluated for EEG seizure detection with an accuracy of 84.2% and for user identification based on walking pattern with an accuracy of 98.4%.
Double Barrier Memristive Devices for Neuromorphic Computing - Martin Zeigler: 2016 International Conference on Rebooting Computing
Martin Zeigler presents his work with fellow collaborators Mirko Hansen and Hermann Kohlstedt, regarding the systems involved with this specific device and its functions.
Read more about this event at: http://icrc.ieee.org/
High Throughput Neural Network based Embedded Streaming Multicore Processors - Tarek Taha: 2016 International Conference on Rebooting Computing
With power consumption becoming a critical processor design issue, specialized architectures for low power processing are becoming popular. Several studies have shown that neural networks can be used for signal processing and pattern recog-nition applications. This study examines the design of memris-tor based multicore neural processors that would be used pri-marily to process data directly from sensors. Additionally, we have examined the design of SRAM based neural processors for the same task. Full system evaluation of the multicore pro-cessors based on these specialized cores were performed taking I/O and routing circuits into consideration. The area and power benefits were compared with traditional multicore RISC pro-cessors. Our results show that the memristor based architec-tures can provide an energy efficiency between three and five orders of magnitude greater than that of RISC processors for the benchmarks examined.
Technology considerations for neuromorphic computing - David Mountain: 2016 International Conference on Rebooting Computing
The use of neural nets has been growing rapidly. A variety of computing architectures, such as CPUs, GPUs, FPGAs, and analog designs have been proposed. This paper will explore how technology ideas affect design choices, using both digital and analog circuit designs suitable for neural nets.
Designing Reconfigurable Large-Scale Deep Learning Systems Using Stochastic Computing - Ao Ren: 2016 International Conference on Rebooting Computing
Deep Learning, as an important branch of machine learning and neural network, is playing an increasingly important role in a number of fields like computer vision, natural language processing, etc. However, large-scale deep learning systems mainly operate in high-performance server clusters, thus restricting the application extensions to personal or mobile devices. The solution proposed in this paper is taking advantage of the fantastic features of stochastic computing methods. Stochastic computing is a type of data representation and processing technique, which uses a binary bit stream to represent a probability number (by counting the number of ones in this bit stream). In the stochastic computing area, some key arithmetic operations such as additions or multiplications can be implemented with very simple components like AND gates or multiplexers, respectively. Thus it provides an immense design space for integrating a large amount of neurons and enabling fully parallel and scalable hardware implementations of large-scale deep learning systems. In this paper, we present a reconfigurable large-scale deep learning system based on stochastic computing technologies, including the design of the neuron, the convolution function, the back-propagation function and some other basic operations.
Neural Processor Design Enabled by Memristor Technology - Hai Li: 2016 International Conference on Rebooting Computing
Deep learning techniques have achieved great suc- cess on various application areas such as image recognition, speech recognition and natural language processing. However, further increasing the scale of deep learning algorithms makes the use of hardware resources fast increasing and becoming unaffordable. Matrix-vector multiplication is a key computing operation in neural processor design and hence greatly affects the execution efficiency. Memristor crossbar is highly attractive for the implementation of matrix-vector multiplication for its analog storage states, high integration density, and built-in parallel execution. In recent years, many memristor based neural processor design have been implement in VLSI hardware system. The current deign schemes can be generally divided into two different approaches: one is referred to "spiking-based" design whose data are represented by digitalized spikes, and the other one is "level-based" design whose data are represented by analog signals with different amplitude. The performance and robustness of the proposed neural process designs are also evaluated by using the application of digital image recognition in previous work. In this work, we investigate the neural processor design that leverages nano-scale memristor technology. A heuristic flow including device modeling, circuit design, architecture, algorithm is studied.
Overcoming the Static Learning Bottleneck - the Need for Adaptive Neural Learning - Craig Vineyard: 2016 International Conference on Rebooting Computing
Amidst the rising impact of machine learning and the popularity of deep neural networks, learning theory is not a solved problem. With the emergence of neuromorphic computing as a means of addressing the von Neumann bottleneck, it is not simply a matter of employing existing algorithms on new hardware technology, but rather richer theory is needed to guide advances. In particular, there is a need for a richer understanding of the role of adaptivity in neural learning to provide a foundation upon which architectures and devices may be built. Modern machine learning algorithms lack adaptive learning, in that they are dominated by a costly training phase after which they no longer learn. The brain on the other hand is continuously learning and provides a basis for which new mathematical theories may be developed to greatly enrich the computational capabilities of learning systems. Game theory provides one alternative mathematical perspective analyzing strategic interactions and as such is well suited to learning theory.
Hyperdimensional Biosignal Processing: A Case Study for EMG-based Hand Gesture Recognition - Abbas Rahimi: 2016 International Conference on Rebooting Computing
The mathematical properties of high-dimensional spaces seem remarkably suited for describing behaviors produces by brains. Brain-inspired hyperdimensional computing (HDC) explores the emulation of cognition by computing with hypervectors as an alternative to computing with numbers. Hypervectors are high-dimensional, holographic, and (pseudo)random with independent and identically distributed (i.i.d.) components. These features provide an opportunity for energy-efficient computing applied to cyberbiological and cybernetic systems. We describe the use of HDC in a smart prosthetic application, namely hand gesture recognition from a stream of Electromyography (EMG) signals. Our algorithm encodes a stream of analog EMG signals that are simultaneously generated from four channels to a single hypervector. The proposed encoding effectively captures spatial and temporal relations across and within the channels to represent a gesture. This HDC encoder achieves a high level of classification accuracy (97.8%) with only1/3 the training data required by state-of-the-art SVM on the same task. HDC exhibits fast and accurate learning explicitly allowing online and continuous learning. We further enhance the encoder to adaptively mitigate the effect of gesture-timing uncertainties across different subjects endogenously; further, the encoder inherently maintains the same accuracy when there is up to 30% overlapping between two consecutive gestures in a classification window.
A Recurrent Crossbar of Memristive Nanodevices Implements Online Novelty Detection - Christopher Bennett: 2016 International Conference on Rebooting Computing
An auto-correlation matrix memory (ACMM) system continuously computes the degree to which a presented input is novel or anomalous relative to past examples. Here we demonstrate that such a filter can be efficiently implemented with memristive nanodevices and accompanying CMOS circuitry. Complete (a full crossbar) and incomplete (an array of memristive devices) variants of the proposed nanofabric are electrically detailed and subsequently simulated on a simple sparse input image test meant to gauge the system's responses to transitions. Both systems demonstrate active novelty filtering with a small level of false positives in the presence of noise, but only the complete system reports all transitions successfully (avoids false negative too). While the system is robust to a noisy channel, degradation towards false positives is more likely when nanodevice variability is taken into account as well. In addition to novelty filtering, the proposed system may be a useful building block for larger reservoir or recurrent on-chip learning systems.
Stochastic Single Flux Quantum Neuromorphic Computing using Magnetically Tunable Josephson Junctions - Stephen Russek: 2016 International Conference on Rebooting Computing
Single flux quantum (SFQ) circuits form a natural neuromorphic technology with SFQ pulses and superconducting transmission lines simulating action potentials and axons, respectively. Here we present a new component, magnetic Josephson junctions, that have a tunablility and re-configurability that was lacking from previous SFQ neuromorphic circuits. The nanoscale magnetic structure acts as a tunable synaptic constituent that modifies the junction critical current. These circuits can operate near the thermal limit where stochastic firing of the neurons is an essential component of the technology. This technology has the ability to create complex neural systems with greater than 10^21 neural firings per second with approximately 1 W dissipation.
Erasing Logic-Memory Boundaries in Superconductor Electronics - Vasili Semenov: 2016 International Conference on Rebooting Computing
Superconducting electronics holds the absolute record for clock frequency and energy efficiency. However, these advantages have so far only given it "a technological back seat" to far more complex CMOS digital circuits. Ironically, the performance and functional density of superconductor circuits decreased when natural finite-state-machine behavior of their RSFQ cells was adapted to memoryless logic gates. We propose to restore this performance by organizing the original "unpasteurized" RSFQ cells into a new family of Memory And loGIC (MAGIC) gate/register objects that run arithmetic calculations as well as store results. The new MAGIC objects eliminate time and energy overheads associated with the conventional transfer of computed data to memory by essentially reducing the transfer distance to zero. The new objects could serve as building blocks for distributed MAGIC-compatible architectures, differing from CMOS-like register files by processing as well as storing data. A simpler Logic Unit (LU) would be sufficient to control the MAGIC registers, because the registers would provide most of the arithmetic functions and separate ALUs would not be needed anymore. The primary goal of the proposed project is to "reboot" the energy consuming process of data exchange between logic and memory.
FinSAL: A Novel FinFET Based Secure Adiabatic Logic for Energy-Efficient and DPA Resistant IoT Devices - Himanshu Thapliyal: 2016 International Conference on Rebooting Computing
With the emergence of Internet of Things (IoT), there is an urgent need to design energy-efficient and secure IoT devices. For example, IoT devices such as Radio Frequency Identification (RFID) tags and Wireless Sensor Nodes (WSN) employ AES cryptographic modules that are susceptible to Differential Power Analysis (DPA) attacks. With the scaling of technology, leakage power in the cryptographic devices increases, which increases the vulnerability to DPA attacks. This paper presents a novel FinFET based Secure Adiabatic Logic (FinSAL), that is energy-efficient and DPA-immune. The proposed adiabatic FinSAL is used to design logic gates such as buffers, XOR, and NAND. Further, the logic gates based on adiabatic FinSAL are used to implement a Positive Polarity Reed Muller (PPRM) architecture based S-box circuit. SPICE simulations at 12.5 MHz show that adiabatic FinSAL S-box circuit saves up to 84% of energy per cycle as compared to the conventional S-box circuit implemented using FinFET. Further, the security of adiabatic FinSAL S-box circuit has been evaluated by performing the DPA attack through SPICE simulations. We proved that the FinSAL S-box circuit is resistant to a DPA attack through a developed DPA attack flow applicable to SPICE simulations.
Neuromorphic computing with integrated photonics and superconductors - Jeffrey Shainline: 2016 International Conference on Rebooting Computing
We present a hardware platform combining integrated photonics with superconducting electronics for large-scale neuromorphic computing. Semiconducting few-photon light-emitting diodes work in conjunction with superconducting-nanowire single-photon detectors to behave as spiking neurons. These neurons are connected through a network of waveguides, and variable weights of connection can be implemented using several approaches. These processing units can operate at $20$ MHz with fully asynchronous activity, light-speed-limited latency, and power densities on the order of 1 mW/cm$^2$. The processing units achieve an energy efficiency of 20 aJ/synapse event, an improvement of roughly a million over recent CMOS demonstrations cite{mear2014}. We present calculations showing this approach could scale to massive interconnectivity near that of the human brain, and could surpass the brain in speed and efficiency.
Cart
Create Account
Sign In